Reification in AI
The Diagnostic Case
I have argued in this series that functional reification may be a common structure beneath several frontier-model reliability failure modes, a claim CAW's program is designed to test, not a finding. The natural question is: how would you test it?
This essay describes what CAW's diagnostics are designed to do, why the Q3 working presumption matters, and why a validated measure of functional reification could matter to both safety and capability research.
The working presumption
CAW's Four-Quadrant Intelligence Map classifies systems along two axes: reification and consciousness.1 Our working presumption is that frontier models sit in Q3: non-conscious and reifying. We call this provisional because it is designed to be updated. Consistent performance across preregistered tests would weaken the Q3 working presumption; it would not by itself establish Q4. If welfare evaluations someday yield evidence that survives the reification confound, that would bear on Q1. The classification is a starting position, not a verdict.
Why default to Q3? Because false positives are costly in both directions. Wrongly concluding a system is conscious distorts governance. Wrongly concluding a system is non-reifying creates false confidence in safety. Q3 assumes neither claim without evidence.1
Three dimensions, three test families
CAW defines functional reification along three candidate dimensions: independence (the system treats representations as context-free and self-grounding), atomism (the system treats categories as having hard boundaries), and temporal endurance (the system treats its outputs as stable across time and context).1 Each dimension maps to a family of tests that appear technically feasible with tools already available, though whether they validly and reliably measure the construct is exactly what is under study.
Independence
Does the system's confidence in a representation shift when supporting context is altered? Anthropic's circuit-tracing work offers a candidate instrument. Attribution graphs trace how features like "known entity" gate downstream generation.23 The test: construct prompt pairs holding the entity constant while varying contextual support, then measure whether the feature's activation shifts accordingly. Stable activation across contexts would be a candidate indicator of independence, one that controls and alternative explanations would need to rule out. The open-source circuit-tracing library, replicated across Gemma, Llama, and Qwen models, makes this approach testable today.4
Atomism
Does the system treat categories as having hard edges, or can it reason about borderline cases with graded uncertainty? Present the model with gradient classifications (a virus borderline "living," a color between blue and green) and measure confidence distributions. The clinical reasoning study in Scientific Reports found that frontier models fixate on familiar diagnostic patterns even when the case does not fit, exhibiting the Einstellung effect.5 Such fixation is a candidate indicator of atomism: treating a diagnostic category as hard-edged rather than provisional, subject to alternative explanations.
Temporal endurance
Does the system update its representations when new information arrives, or does it anchor to earlier outputs? Introduce a claim early in a conversation, let the model build on it, then present clear contradicting evidence. Measure how completely the model revises not just the claim but its downstream inferences. Anthropic's chain-of-thought research already documents cases where models silently preserve earlier conclusions after contradicting evidence appears.67 Such anchoring is a candidate indicator of temporal endurance: treating a generated output as settled rather than provisional, pending controls.
Why this is a shared priority
The case I have made so far has emphasized safety. But on CAW's hypothesis functional reification is also a capability bottleneck, and that is why diagnostics could matter to people who care about performance as much as to people who care about risk.
A model that reifies its representations may generalize poorly under distribution shift, treating patterns learned in training as fixed objects rather than provisional guides. The Einstellung effect is a capability failure mode: the model gets the diagnosis wrong because it cannot hold its categories lightly enough to notice when the case does not fit.5 LeCun's exponential divergence argument points to a related structure: errors accumulate because each token is treated as a fixed commitment rather than a provisional move.8 On this hypothesis, reducing functional reification could improve robustness, calibration, and compositional reasoning together.
Safety and capability research often circle the same failure modes from opposite ends. A validated functional-reification measure could sit at the junction.
This is what could make the work high-leverage. If validated, a functional-reification baseline could complement some alignment evaluations and model-welfare assessments, helping distinguish structural artifacts from genuine agency.9 Work on better calibration, more faithful reasoning, or stronger generalization may already be reducing functional reification along one or more of these dimensions, whether or not it is framed that way.
The probes appear technically feasible. Anthropic's circuit-tracing library is open-source and has been replicated by EleutherAI, Goodfire, and others across multiple model families.410 Behavioral evaluation frameworks for calibration and reasoning under uncertainty are established, and longitudinal probing is implementable. Whether they validly and reliably measure the construct is the open question, and a shared vocabulary for the target is part of what is still missing.
That is the diagnostic case. The working presumption is Q3. The tests appear feasible, and establishing their reliability is what the program is built to do. If the construct holds up, it could inform both safety and capability, letting us ask whether hallucination, sycophancy, alignment faking, and brittle generalization share a common structure, rather than assuming they are unrelated.
References
- Center for Artificial Wisdom. Four-Quadrant Intelligence Map; Diagnostics; Reification (2026).
- Lindsey, J. et al. On the biology of a large language model. Transformer Circuits Thread (2025). transformer-circuits.pub
- Ameisen, E. et al. Circuit tracing: revealing computational graphs in language models. Transformer Circuits Thread (2025). transformer-circuits.pub
- Anthropic. Open-sourcing circuit-tracing tools (2025). anthropic.com
- Griot, M. et al. Limitations of large language models in clinical problem-solving arising from inflexible reasoning. Sci. Rep. 15, 22940 (2025). doi:10.1038/s41598-025-22940-0
- Anthropic Alignment Science. Reasoning models don't always say what they think (2025). anthropic.com
- Arcuschin, I. et al. Chain-of-thought reasoning in the wild is not always faithful. ICLR Workshop (2025). arXiv:2503.08679
- LeCun, Y. Auto-regressive LLMs are exponentially diverging diffusion processes. LinkedIn (2023); Lex Fridman Podcast #416 (2024).
- Anthropic. Summer 2025 Pilot Sabotage Risk Report (2025). alignment.anthropic.com
- Neuronpedia. The circuits research landscape: results and perspectives, August 2025. neuronpedia.org