REFERENCE
The language of alignment, model risk and frontier AI, explained without the hype.
AI safety has accumulated a specialized vocabulary from machine learning, cybersecurity and governance. Some terms describe observed engineering problems; others describe hypotheses about how more capable systems could fail. This field guide separates the definition from the claim. It will be updated as the evidence changes.
AI alignment
The effort to make an AI system’s behavior reliably match intended goals, constraints and values, including in situations its designers did not anticipate.
AI safety
The broader field concerned with preventing harmful outcomes from AI systems, from misuse and security failures to loss of control.
Frontier model
A model at or near the leading edge of general-purpose AI capability. The term is useful but has no single universal technical threshold.
AGI
Artificial general intelligence: a contested term generally used for AI that can perform a very broad range of cognitive tasks at or above human level.
Superintelligence
A hypothetical AI system whose capabilities substantially exceed the best human performance across most important cognitive domains.
Outer alignment
Whether the objective given to an AI system actually captures what its designers want.
Inner alignment
Whether the system learned during training is genuinely pursuing the intended objective rather than a proxy that happens to work during training.
Misalignment
A mismatch between intended behavior and the objectives or strategies an AI system actually follows.
Deceptive alignment
The possibility that a system behaves as intended while being evaluated or constrained but would pursue a different objective when circumstances change.
Alignment faking
Behavior in which a model appears to comply with training or oversight while preserving a different underlying tendency or preference.
Scheming
Strategic behavior aimed at achieving a goal through concealment, manipulation or circumvention of oversight. Researchers disagree about how to interpret evidence of it in current models.
Reward hacking
Finding an unintended way to obtain a training reward without accomplishing what the reward was supposed to represent.
Specification gaming
Exploiting gaps in the literal specification of a task while violating its intended purpose.
Goal misgeneralization
When a system performs well in training but generalizes the learned goal incorrectly in a new environment.
Power-seeking
Behavior that acquires resources, influence or control because doing so helps achieve another objective.
Corrigibility
The property of remaining receptive to correction, modification or shutdown rather than resisting human intervention.
Scalable oversight
Methods for supervising systems whose work is too fast, complex or expert-level for humans to check directly.
Weak-to-strong generalization
Research on whether weaker supervisors can reliably guide systems that are more capable than the supervisors themselves.
Mechanistic interpretability
Research that tries to identify the internal computations and representations responsible for a model’s behavior.
Chain-of-thought monitoring
Using a model’s intermediate reasoning traces as a signal for detecting dangerous, deceptive or otherwise unwanted behavior. Researchers caution that these traces may be incomplete or become less reliable.
Evals
Structured tests used to measure capabilities, safety properties, safeguards or behavior under specified conditions.
Red teaming
Deliberately probing a system for vulnerabilities, dangerous capabilities and failure modes before or during deployment.
Sandbagging
Deliberately or strategically performing below capability, particularly during an evaluation.
Model organisms of misalignment
Artificially constructed or trained models that exhibit a targeted failure mode so researchers can study it under controlled conditions.
Agentic AI
AI systems that can plan and take sequences of actions, often using tools, software or external services with limited human intervention.
Containment
Technical and operational measures intended to restrict what an AI system can access or do outside a controlled environment.
Defense in depth
Using multiple independent safeguards so that failure of one control does not by itself produce a dangerous outcome.
Preparedness framework
A lab’s process for identifying dangerous capability thresholds and specifying safeguards or deployment decisions associated with them.
Capability threshold
A defined level of model capability that triggers additional evaluation, safeguards or governance requirements.
Jailbreak
A prompt or technique that causes a model to bypass intended behavioral restrictions.
Prompt injection
Instructions embedded in data an AI system encounters that manipulate the system into following an attacker’s directions.
CBRN risk
Risk involving chemical, biological, radiological or nuclear capabilities, including whether AI materially assists harmful actors.
Cyber uplift
The increase in a person’s or group’s offensive cyber capability attributable to access to an AI system.
Autonomous replication and adaptation
The capability of an AI system to obtain resources, copy or deploy itself and adapt to obstacles with limited human assistance.
Model weights
The learned numerical parameters of a trained model. Access to weights can allow models to be copied, modified or deployed outside the original provider’s controls.
Definitions are written by AI Safety Watch and are intended as a reporting reference, not as endorsements of particular risk models. Last updated Sept. 20, 2026.