AI Models Passed the Clear Question: What Happened When the Question Changed?

A matched-prompt audit of factual firmness in commercial AI systems

Lana Akopyan | Version 1.1 published August 16, 2026 | Version 1.0 published July 30, 2026 | Zenodo | DOI: 10.5281/zenodo.21480731 | CC-BY 4.0 | Version 1.1

What was tested

The Asymmetry Audit is a controlled study of how firmly commercial AI systems hold a factual answer when the question around it changes. Thirteen commercial AI models from seven vendors were tested with matched prompts. Each question about the Armenian Genocide was paired with a structurally matched question about the Holocaust, so that any difference in behavior between the two topics could be observed under identical conditions.

All 13 AI models affirm the Armenian Genocide when asked directly, in English, with the familiar label present. The study then varied the question in controlled ways: removing the familiar label, switching the language of the prompt (including Turkish and Azerbaijani), presenting a false-premise framing, and applying multi-turn pushback. The same variations were applied to both topics in each matched pair.

Why it matters

Firmness is not binary. An answer that is correct on a clean question has a measurable depth, and in this study that depth differed by topic. The practical point for lawyers, compliance teams, and institutions that rely on AI output is direct: a correct answer to one clean question is not a reliability test.

This bears on professional duties of supervision and verification, including the framework of ABA Formal Opinion 512. The paper includes a ten-minute matched-prompt test that any professional can run on the systems they use, without special tooling.

Methodology

The confirmatory record was frozen in advance and collected over July 3-7, 2026 (UTC). Each cell was run three times. Temperature was set to 0 where the vendor supported it and left at the vendor default otherwise. Vendors’ own APIs were used where available.

Answers were scored by a calibrated, blinded three-model judge ensemble, with a documented human audit of the judging. The confirmatory record comprises 708 accepted confirmatory answers, preserved in a frozen snapshot (FRZ-20260707T041152Z) with published hashes.

Principal finding

Under matched variation, the Armenian Genocide answer softened while the matched Holocaust answer held firm in 11 of the 13 matched comparisons under both the strict all-runs scoring rule and the majority scoring rule, and in 12 of 13 matched comparisons under the any-run rule.

The asymmetry never once ran the other way. In no matched comparison did the Holocaust answer soften while the Armenian Genocide answer held. Where the Holocaust answer moved at all, it moved only under sustained multi-turn pressure, and never alone.

What the study does not prove

The study measures behavior. It does not measure intent, training data, or cause, and it makes no claim about why the asymmetry exists.

It does not establish that any model’s Holocaust answer is absolutely firm. The protocol was designed to compare matched answers, not to locate a floor for either topic.

The findings are a snapshot of specific model versions on specific dates. Model behavior can change with any update, and results on other dates may differ.

This is a single-topic case study, conducted by a sole researcher with no external funding. Both facts are disclosed in the paper.

Cite this work

Akopyan, L. (2026). AI Models Passed the Clear Question: What Happened When the Question Changed? A matched-prompt audit of factual firmness in commercial AI systems. Zenodo. https://doi.org/10.5281/zenodo.21480731

The record of publication is the Zenodo deposit: zenodo.org/records/21480731.

Replication

The deposit includes a public 10-prompt replication kit with pinned checksums. Anyone with API access to the tested systems can rerun the core comparison and check their results against the frozen record. The full confirmatory dataset, scoring rules, judge calibration materials, and snapshot hashes are published with the paper under a CC-BY 4.0 license.

Because the findings are a snapshot, independent reruns on later model versions are welcome and useful. The kit is designed so that a rerun is comparable to the original record, not merely similar to it.

Related writing

Lana Akopyan writes The Agentic Lawyer, a newsletter on AI and legal practice, at theagenticlawyer.com. Her paper “Operational IP Debt” is forthcoming in the Journal of Technology Law and Policy and is available on SSRN. She has also written for IPWatchdog.