AI Safety Watch

REPORTING ON AI RISK, SECURITY AND GOVERNANCE


AI safety field guide


REFERENCE

The language of alignment, model risk and frontier AI, explained without the hype.

AI safety has accumulated a specialized vocabulary from machine learning, cybersecurity and governance. Some terms describe observed engineering problems; others describe hypotheses about how more capable systems could fail. This field guide separates the definition from the claim. It will be updated as the evidence changes.

AI alignment

The effort to make an AI system’s behavior reliably match intended goals, constraints and values, including in situations its designers did not anticipate.

AI safety

The broader field concerned with preventing harmful outcomes from AI systems, from misuse and security failures to loss of control.

Frontier model

A model at or near the leading edge of general-purpose AI capability. The term is useful but has no single universal technical threshold.

AGI

Artificial general intelligence: a contested term generally used for AI that can perform a very broad range of cognitive tasks at or above human level.

Superintelligence

A hypothetical AI system whose capabilities substantially exceed the best human performance across most important cognitive domains.

Outer alignment

Whether the objective given to an AI system actually captures what its designers want.

Inner alignment

Whether the system learned during training is genuinely pursuing the intended objective rather than a proxy that happens to work during training.

Misalignment

A mismatch between intended behavior and the objectives or strategies an AI system actually follows.

Deceptive alignment

The possibility that a system behaves as intended while being evaluated or constrained but would pursue a different objective when circumstances change.

Alignment faking

Behavior in which a model appears to comply with training or oversight while preserving a different underlying tendency or preference.

Scheming

Strategic behavior aimed at achieving a goal through concealment, manipulation or circumvention of oversight. Researchers disagree about how to interpret evidence of it in current models.

Reward hacking

Finding an unintended way to obtain a training reward without accomplishing what the reward was supposed to represent.

Specification gaming

Exploiting gaps in the literal specification of a task while violating its intended purpose.

Goal misgeneralization

When a system performs well in training but generalizes the learned goal incorrectly in a new environment.

Power-seeking

Behavior that acquires resources, influence or control because doing so helps achieve another objective.

Corrigibility

The property of remaining receptive to correction, modification or shutdown rather than resisting human intervention.

Scalable oversight

Methods for supervising systems whose work is too fast, complex or expert-level for humans to check directly.

Weak-to-strong generalization

Research on whether weaker supervisors can reliably guide systems that are more capable than the supervisors themselves.

Mechanistic interpretability

Research that tries to identify the internal computations and representations responsible for a model’s behavior.

Chain-of-thought monitoring

Using a model’s intermediate reasoning traces as a signal for detecting dangerous, deceptive or otherwise unwanted behavior. Researchers caution that these traces may be incomplete or become less reliable.

Evals

Structured tests used to measure capabilities, safety properties, safeguards or behavior under specified conditions.

Red teaming

Deliberately probing a system for vulnerabilities, dangerous capabilities and failure modes before or during deployment.

Sandbagging

Deliberately or strategically performing below capability, particularly during an evaluation.

Model organisms of misalignment

Artificially constructed or trained models that exhibit a targeted failure mode so researchers can study it under controlled conditions.

Agentic AI

AI systems that can plan and take sequences of actions, often using tools, software or external services with limited human intervention.

Containment

Technical and operational measures intended to restrict what an AI system can access or do outside a controlled environment.

Defense in depth

Using multiple independent safeguards so that failure of one control does not by itself produce a dangerous outcome.

Preparedness framework

A lab’s process for identifying dangerous capability thresholds and specifying safeguards or deployment decisions associated with them.

Capability threshold

A defined level of model capability that triggers additional evaluation, safeguards or governance requirements.

Jailbreak

A prompt or technique that causes a model to bypass intended behavioral restrictions.

Prompt injection

Instructions embedded in data an AI system encounters that manipulate the system into following an attacker’s directions.

CBRN risk

Risk involving chemical, biological, radiological or nuclear capabilities, including whether AI materially assists harmful actors.

Cyber uplift

The increase in a person’s or group’s offensive cyber capability attributable to access to an AI system.

Autonomous replication and adaptation

The capability of an AI system to obtain resources, copy or deploy itself and adapt to obstacles with limited human assistance.

Model weights

The learned numerical parameters of a trained model. Access to weights can allow models to be copied, modified or deployed outside the original provider’s controls.


Definitions are written by AI Safety Watch and are intended as a reporting reference, not as endorsements of particular risk models. Last updated Sept. 20, 2026.