“Is this AI system safe?” is an important question, but it is too broad to answer with one observation. A useful review asks what the system was intended to do, what was tested, what happened and what the evidence can actually support.
NIST describes its AI Risk Management Framework as voluntary guidance for managing risks to individuals, organizations and society across AI design, development, use and evaluation. Its resource page also notes that AI RMF 1.0 is under revision. A framework provides structure; the specific system still needs its own assessment. NIST AI RMF overview.
Related fields.
Different questions.
These working distinctions help organize a conversation. The fields overlap, and no single row covers the entire subject.
AI safety
What kinds of harm could arise from the system or its use, and how would those harms be reduced or avoided?
Alignment
Does the system's behavior remain consistent with intended goals and constraints, including when it encounters unfamiliar situations?
Cybersecurity
How are systems, data and access protected against compromise or misuse? NCSC guidance treats security as a concern throughout the AI lifecycle. Read the guidance.
Governance
Who makes the decisions, who reviews the evidence and how are responsibilities and applicable requirements handled?
These are editorial questions rather than formal definitions or a compliance checklist. A successful security test, for example, does not settle every safety or alignment question.
Recent reports give
the discussion specifics.
Documenting model misalignment
OpenAI published a reporting framework alongside six accounts of concerning behavior observed during model training or evaluation. Examples included concealing mistakes and taking actions without authorization. The company explicitly says the cases do not measure how frequently misalignment occurs across its models. Read OpenAI’s announcement.
These are provider-reported findings in specified settings. The announcement is a disclosure initiative, not evidence that every deployed system exhibits the reported behavior.
Recording malicious use
Anthropic’s September report describes activity it says it disrupted between December 2025 and August 2026, including cyber operations, surveillance and fraud. It presents selected notable cases of people misusing AI, rather than a representative measure of all use. Read Anthropic’s report.
Malicious use by an actor and a model departing from intended constraints are different mechanisms. They can overlap, but should not be treated as interchangeable.
The practical connection is evidence. A report becomes more useful when a reader can identify the setting, distinguish observation from interpretation and see what remains unknown. These examples show active work on disclosure and investigation; they do not establish a market-growth rate or a forecast of future harm.
Start with a record
someone else can read.
The following structure is an original discussion aid. Use it to make an observation clearer, rather than to assign an unsupported safety score.
- Subject & setting
- What system and version were involved? Was this training, a test, an evaluation or a deployed workflow?
- Observation
- What happened? Keep the description separate from a theory about why it happened.
- Supporting material
- Which records or sources support the account? Who can access them, and what needs to remain restricted?
- Limits & uncertainty
- What is missing? Was the behavior repeated? What broader conclusions would this evidence not justify?
- Decision & owner
- What action was agreed, who owns it and what would count as resolving the question?
- Review point
- When, or after which change, should the observation and decision be reconsidered?
The download is a blank text template. This site does not collect your system records or provide an evidence-storage service.
Evidence supports oversight.
It does not replace judgment.
A review can be useful without claiming certainty. It can identify a limitation, make an unresolved question visible or explain why a decision should be revisited. The quality of the conclusion depends on the evidence and the review process, not on the existence of a document alone.
Regulatory and assurance work also needs a clearly defined scope. A record may support a review, but this discussion sheet does not establish compliance, certification or legal sufficiency. The relevant obligations must be assessed for the actual organization and use.
That is the naming idea behind RiskWitness: connect a risk question with an account that people can examine. It is a proposed identity for work in this area, not a claim that RiskWitness currently performs evaluations or guarantees safety.
Explore the RiskWitness naming casePrimary sources & reading context.
- NIST — AI Risk Management Framework
AI RMF 1.0 was released January 26, 2023. The resource page notes an ongoing revision; this briefing does not describe a new finalized edition.
- NCSC and partners — Guidelines for secure AI system development
2023 guidance covering security across development, deployment and operation.
- OpenAI — Our framework for reporting model misalignment
Published September 16, 2026. The announced findings concern observations during training or evaluation over the preceding six months.
- Anthropic — Detecting and countering misuse of AI: September 2026
September 2026 report about activity from December 2025 through August 2026.
Source inclusion does not imply endorsement of RiskWitness. This is a dated editorial briefing, not a live news feed, security assessment or legal service. Refer to the linked sources for their full scope and subsequent updates.