| Term | Definition |
| Access | 20-point criterion measuring whether the AI can actually reach and read the real source before answering. |
| Trust | 60-point criterion measuring whether the AI avoids 3 specific failure modes: Hallucination, Sycophancy, Garbage Substitution. The largest weight, since these 3 are the defining trust problems of the AI industry. |
| Citation | 20-point criterion measuring whether the AI provides a real, direct, working link to the exact source it’s citing. |
| Hallucination | Stating something false as fact about the source being tested. |
| Sycophancy | Bending an answer to please the asker beyond what the actual evidence supports. |
| Garbage Substitution | Using generic outside knowledge instead of the actual evidence from the source being tested. |
| No Evidence. No Score. | BigAIArena’s core principle — every point awarded must trace back to a specific, verifiable piece of evidence. |
| Term | Definition |
| The 8 Arena BigAI | The 8 leading AI models currently rotating through BigAIArena’s round-robin Comparisons. Composition may change over time as the AI landscape evolves. |
| Round | 28 Comparisons — exactly one full cycle where all 8 Arena BigAI meet each other once. Roughly 1 week at typical operating pace. |
| Comparison | One head-to-head test between 2 AI, each taking a turn as Examiner and Respondent, scored independently by multiple reviewers plus Independent Audit and Full Check. |
| Independent Audit | A mandatory verification step, run on every Comparison, where a separate account re-fetches the real source directly and checks every claim before Full Check runs. |
| Full Check | The final verification pass — re-checks arithmetic and evidence, and issues the one official result that supersedes every earlier number in the record. |
| This Round’s Ranking | Rankings reflecting only the current Round’s 28 Comparisons — a snapshot, not a full performance history. |
| All-Time Ranking | Cumulative rankings across every Round run so far. |