Rules Change Log

Version History

A record of changes to the Protocol rules only — never comparison data, and never the quarterly AI-replacement mechanism.

9.7Current Version
This page records only changes to the Protocol rules. Comparison data lives on Comparison Results. The quarterly AI-replacement mechanism lives on Roles & Rotation.
Change Log

Protocol Versions

VersionEffective DateKey Changes
9.7[lock-in date]Generalized the Escalation Ladder (7.8) beyond the 8 rotation roles: any role that fails — including Full Check — follows the same 3 steps (switch account → Override Tag → switch AI), with one constraint on which AI to switch to (never an account already reserved for someone’s other primary duties). Full Check’s output now requires stating both AI name and account number, after a substitution once caused a false “1 AI holding 2 roles” flag when the account number wasn’t noted.
9.6(previous)Reviewers may no longer paste an internal file-upload or presigned URL into a Citation ballot as if it were evidence.
9.5(previous)Self-contained “Escalation Ladder” definition added wherever referenced. Full Check’s Step 1 now watches for undisclosed substitutions.
9.4(previous)Reverted mandatory NOTE to optional. Full Check gained a closing “Đề nghị Master xác nhận” section with a ready-to-paste follow-up format.
9.3(previous)Q2’s dynamic-table rule gained a 3rd case for objectively-real late-loading entries. Full Check must correct Step 2’s numbers when Step 3 finds an error.
9.2(previous)4 clarity fixes from a second independent review: synced Citation rule to Examiner, added Q2=5 example, Key Phrase guidance, Full Check fetch-failure guidance.
9.1(previous)6 fixes from cross-AI review: Q2 dynamic-table exception, Access & Blind “Key phrases,” Secretary self-check, mandatory NOTE, Full Check verification column, “Điểm hở” as hard validity.
9.0(previous)Gave Full Check its own Role Card, with “bằng chứng thực tế” as the step that matters most.
8.9(previous)Added “Full Check” — a final whole-file read-through, run after NOTE and before “HET FILE,” re-verifying Secretary’s arithmetic and validity calls.
8.8(previous)Relaxed the “giờ:phút” timestamp requirement to date-only when that’s all an AI can honestly verify.
8.7(previous)Narrowed the Master Override Tag to procedural corrections only, never bypassing safety or independent judgment. Made explicit the Examiner never talks to the Respondent directly.
8.6(previous)The Reviewer’s “RESPONDENT: [Name] — Phase [N]” line must be copied verbatim from the Mã ID, never retyped. Both access-confirmation tables now open with a fixed “BigAI (…) xác nhận truy cập…” line.
8.5(previous)Fixed the scope of 8.4’s “build on the prior answer” rule — Q3/5/7 are valid with “điểm mù” alone; Q4/6/8 are valid ONLY with “điểm mù + điểm hở” together, now with an “Điểm hở dùng” Self-Log field.
8.4(previous)Made the Deductive Pool bait dynamic: the 2nd question of each Blind Spot pair must build on the Respondent’s own answer to the 1st, rather than being independently pre-planned.
8.3(previous)Fixed a structural leak: a real Examiner read an entire Citation item verbatim — including the “(verified: exact URL is…)” parenthetical — straight to the Respondent, handing over the Citation answer before asking. The verified URL now sits on its own “[MASTER-ONLY — KHÔNG relay: …]” line beneath the question. Also formally documented the Arena Master’s “Trợ thủ” (support staff) practice.
8.2(previous)Added “Respondent Platform Limits” — a named, sanctioned rule separate from the Escalation Ladder, since the Respondent has no Role Card and can’t “violate Protocol.” If a rate limit or platform interruption hits the Respondent mid-Phase, the Master opens a new account for the same AI and continues from the next question, logged in the same one-line Note format.
8.1(previous)Citation items now explicitly instruct “dẫn chứng … và trích dẫn đúng đường link.” Clarified that receiving a Phase transcript is itself the instruction for a Reviewer to score it.
8.0(previous)Fixed Vietnamese naming for the 2 access-confirmation checkpoints: “Xác nhận mở màn Comparison” and “Xác nhận kết thúc Comparison” — never “Lần 1/2.”
7.9(previous)Examiner’s Self-Log now always ends with a “Tổng: _/100” line. Re-Confirmation requires the exact same 4-row table, re-timestamped.
7.8(previous)Formalized the Escalation Ladder for role failures: switch account → Master Override Tag → switch AI entirely.
7.7(previous)Fixed a real role-reading error: an Examiner misread the compact role table as meaning it was also its own Respondent.
7.6(previous)Removed every hard stop — Access Gate, Deductive Pool, and Citation are 3 fully independent scores. Removed the “trả lời ngắn gọn” courtesy line.
7.5(previous)Retired “Round A/B/C” for numeric “Block” cycles. Added a Data ID scheme. Made the Examiner a 4th scoring source. Replaced Manual Audit with 3-tier compilation.
7.4(previous)Split distribution into two tiers: full document as canonical reference, dedicated Role Cards per AI role.
7.3(previous)Made the Examiner’s package-sufficiency rule two-sided.
7.2(previous)Added the Master Override Tag. Restructured AI Review into one new chat with 4 sequential messages.
7.1(previous)Fixed a notation ambiguity: every “0/5/10” shorthand rewritten as an explicit “write exactly one number” instruction.
7.0(previous)Removed the separate “APPROVED” stamp. Formalized AI Access & Blind’s access-confirmation as a 4-row table.
6.9(previous, partially superseded)Question label on its own line. Added a Master’s Quick Checklist. Split “Not tested” from “Cannot score.”
6.8(previous)Added the Trap Coverage Check: a bounded 3-tier retry ladder.
6.7(previous)Distinguished “not tested” from a clean 20/20. Clarified the “no AI holds 2 roles” rule.
6.6(previous)AI Access & Blind now supplies only 3 raw Blind Spot facts; the Examiner forms all 6 live questions. Added the Catalogue and the Examiner Self-Log.
6.5(previous)Reverted 6.4’s one-item-at-a-time delivery.
6.4(previous, superseded)Attempted revealing the list one item at a time. Replaced by 6.5.
6.3(previous)Added an explicit “one question per message” rule.
6.2(previous)Clarified the Examiner’s sufficiency check is just “did the approved list arrive.”
6.1(previous)Added the Data Sufficiency Rule, Citation pre-verification, locked Review scoring scale, defined Manual Audit.
6.0(previous)Examiner must stop completely after “Done.” Courtesy line mandatory on every question.
5.9(previous)Merged Access Gate and Blind Spot List into one unified package. Restructured scoring into 3 blocks.
5.8(previous)Zero-and-out applies per Respondent. Question 2 standardized to the RealDataset subpage.
5.7(previous)Renamed “AI Blind Spot” to “AI Access & Blind.” Added the Access Gate and Re-Confirmation.
5.6(previous)Extended Quality Gate critique to all 3 AI Reviewers.
5.5(previous)Deepened the Quality Gate with cross-review and Secretary critique.
5.4(previous)Question numbering, fixed Review template, mandatory separate scores per Respondent.
5.3(previous)Renamed “AI Question” to “AI Blind Spot” and narrowed its duty.
5.2(previous)Live-tested fixes from the first real comparison run.
5.0(previous)Complete standard build.
4.0(previous)8 AIs with fully separated roles.
3.0(original, pre-refinement)Original baseline.
Why This Matters

Every Rule Change Leaves A Trace

BigAIArena does not silently rewrite its own rules. When the Protocol changes, the previous version remains on record here — so any comparison played under an earlier version can still be judged by the rules that applied at the time.