Skip to content

Research

AI Incident Disclosure Baseline (AIDB) v1.0

A rigorous, implementable baseline for identifying, classifying, investigating, and disclosing AI incidents across technical, legal, and societal risk domains.

governancepolicydisclosureincident-responseai-safety

7 September 2026 · Reseni Governance Team

AI systems can fail without crashing. They can produce harmful recommendations while remaining available, expose memorised data through ordinary interfaces, discriminate at scale without a single anomalous transaction, or enable misuse through capabilities that were neither intended nor adequately constrained. Conventional security incident processes capture part of this problem, but not all of it.

The AI Incident Disclosure Baseline (AIDB) establishes a minimum standard for deciding when an AI-related event becomes a reportable incident, what evidence an organisation should preserve, who should receive notice, and what a defensible disclosure must contain. It is designed for organisations that develop, provide, deploy, integrate, or materially depend on AI systems.

AIDB is a voluntary baseline, not a substitute for applicable law, sector regulation, contractual duties, or emergency reporting obligations. Its purpose is to create a common operational floor where no globally adopted disclosure protocol yet covers the full range of AI harms.

Executive position

An organisation should disclose an AI incident when there is credible evidence that an AI system has caused, materially contributed to, or created an imminent and plausible risk of significant harm, and disclosure would enable affected parties, operators, regulators, or the public to reduce that harm.

The decision must not depend solely on whether:

  • malicious exploitation occurred;
  • a software vulnerability was assigned a CVE;
  • personal data was exfiltrated;
  • the model behaved deterministically;
  • the organisation intended the affected use;
  • the incident crossed a financial materiality threshold; or
  • root cause has been conclusively established.

AI incident disclosure is therefore broader than vulnerability disclosure. It includes failures of safety, security, privacy, fairness, reliability, transparency, human oversight, and lawful operation.

1. Scope

AIDB applies across the AI lifecycle:

  1. data collection, acquisition, labelling, and preparation;
  2. model design, training, fine-tuning, and evaluation;
  3. system integration, retrieval, tool use, and agent orchestration;
  4. deployment, monitoring, and human oversight;
  5. model or service updates;
  6. retirement, withdrawal, and downstream decommissioning.

It covers both internally developed and third-party AI, including foundation models, embedded models, decision-support systems, autonomous or semi-autonomous agents, recommender systems, biometric systems, synthetic-media systems, and AI-enabled security controls.

The accountable organisation is not limited to the model developer. Providers, deployers, integrators, distributors, data suppliers, and operators may each hold evidence or control mitigations that no other actor possesses.

2. Normative language

The terms MUST, MUST NOT, SHOULD, SHOULD NOT, and MAY indicate the strength of a requirement within this voluntary baseline:

  • MUST: necessary to claim conformance;
  • SHOULD: expected unless a documented risk-based justification supports an alternative;
  • MAY: optional and context-dependent.

Conformance does not establish legal compliance or certify that an AI system is safe.

3. Core definitions

AI incident

An event, series of events, or continuing condition in which the development, deployment, or use of an AI system causes, contributes to, or materially increases the likelihood of harm to people, property, organisations, critical infrastructure, democratic processes, the environment, or fundamental rights.

An incident may arise from:

  • expected use;
  • reasonably foreseeable misuse;
  • malicious use or compromise;
  • interaction with another system;
  • deficient human oversight;
  • data, model, interface, or policy changes;
  • scale effects not visible during pre-deployment testing; or
  • a mismatch between represented and actual system capabilities.

AI hazard

A condition, capability, behaviour, or failure mode that could lead to an incident but has not yet produced verified harm. A hazard becomes reportable under this baseline when exposure is material, exploitation is credible, or preventive action by others is time-sensitive.

Near miss

An event that could reasonably have produced material harm but did not, because of chance, timely intervention, limited exposure, or an effective control. Near misses MUST be recorded internally and SHOULD be disclosed when they reveal a systemic risk shared by downstream users or comparable systems.

AI vulnerability

A weakness in an AI system or its operational context that can be exploited or triggered to violate a security, safety, privacy, fairness, reliability, or governance objective. Vulnerabilities include, but are not limited to, prompt injection, unsafe tool invocation, model extraction, training-data leakage, evaluation bypass, poisoned retrieval sources, and failures in access or policy enforcement.

Material harm

Harm is material when its severity, scale, duration, reversibility, distribution, or effect on protected interests is significant enough that a reasonable affected party or accountable authority would need the information to make a protective decision.

4. Disclosure principles

4.1 Harm reduction

Disclosure exists to reduce harm, not merely to document it. Timing, content, and distribution MUST be selected according to what recipients need to protect themselves without creating disproportionate secondary risk.

4.2 Evidence before certainty

Organisations MUST distinguish confirmed facts, credible assessments, hypotheses, and unknowns. They MUST NOT delay an initial notice solely because causal analysis is incomplete.

4.3 Proportional transparency

The detail disclosed SHOULD increase with the incident's severity, public impact, and downstream relevance. Sensitive exploit instructions, personal data, protected security information, and details that create an immediate misuse pathway MAY be withheld, but the reason and review date SHOULD be recorded.

4.4 Human-centred materiality

Severity MUST be assessed from the perspective of affected people and institutions, not only from the provider's financial, technical, or reputational exposure.

4.5 Non-retaliation

Organisations SHOULD maintain a good-faith reporting policy covering employees, users, researchers, contractors, and downstream operators. Contractual terms MUST NOT be used to suppress lawful reporting of credible harm.

4.6 Correctability

Every disclosure MUST be versioned. Material corrections, scope changes, and newly identified affected populations MUST be published through the same channels as the original notice.

4.7 Accessibility

Notices intended for affected people MUST use plain language, accessible formats, and translations appropriate to the population at risk. A technical appendix MAY accompany, but MUST NOT replace, an actionable public explanation.

5. Incident classification

An organisation MUST evaluate severity across at least six dimensions:

DimensionAssessment question
ImpactWhat physical, psychological, economic, legal, privacy, security, or rights-based harm occurred or could occur?
ScaleHow many people, systems, organisations, jurisdictions, or decisions are affected?
ExposureIs the behaviour isolated, reproducible, actively exploited, or structurally present across deployments?
ReversibilityCan affected parties be restored, compensated, notified, or protected before consequences become permanent?
VulnerabilityAre children, patients, workers, migrants, dissidents, or other at-risk groups disproportionately exposed?
Systemic reachCan the failure propagate through model supply chains, shared APIs, critical infrastructure, or high-impact decisions?

The highest credible dimension governs the initial severity. An organisation MUST NOT average away a catastrophic impact because the observed case count is low.

Severity levels

LevelClassificationMinimum thresholdIllustrative cases
AIDB-1CriticalActual or imminent severe harm with broad, irreversible, life-safety, critical-infrastructure, national-security, or systemic implicationsAI-enabled control failure affecting essential services; scalable capability enabling severe misuse; widespread unlawful biometric identification
AIDB-2HighSerious harm, significant rights impact, sensitive-data exposure, active exploitation, or repeated harmful decisions affecting a defined populationmaterial clinical decision error; exploitable cross-tenant data leakage; systematic discrimination in employment or credit
AIDB-3ModerateLimited or reversible harm requiring corrective action, notification, or downstream mitigationbounded harmful recommendation pattern; material transparency failure; recurring unsafe output under foreseeable use
AIDB-4Low / learning eventNo verified material harm and low immediate exposure, but evidence is relevant to prevention, assurance, or trend analysiscontained evaluation failure; near miss blocked by human review; non-sensitive benchmark regression

Examples are non-exhaustive. Sector-specific obligations may classify the same event more severely.

6. Reporting clocks

The clock begins when an authorised incident function has, or reasonably should have, credible evidence that the reporting threshold may be met. It does not begin only after executive confirmation or completion of root-cause analysis.

The following are baseline targets, not representations of statutory deadlines:

SeverityInternal escalationInitial external noticeSubstantive updateFinal or stable report
AIDB-1Immediate; target within 1 hourWithout undue delay; target within 24 hoursAt least every 24 hours while acute risk continuesTarget within 30 days
AIDB-2Target within 4 hoursWithout undue delay; target within 72 hoursAt least every 7 days while unresolvedTarget within 45 days
AIDB-3Target within 1 business dayTarget within 7 calendar days when external action is neededOn material changeTarget within 60 days
AIDB-4Normal risk workflowExternal notice optional unless aggregation reveals material riskAs appropriateInternal closure record

Where law, regulation, contract, or sector policy imposes a shorter period, the shorter period controls. Where premature public disclosure would materially increase exploitation or safety risk, an organisation MAY use coordinated disclosure, but MUST still notify parties capable of reducing immediate harm.

7. Recipient and routing model

Disclosure is not a single press release. The incident lead MUST identify recipients according to the actions they can take.

Affected people

Notify affected people when they need to:

  • stop relying on an AI-generated decision or output;
  • seek medical, legal, financial, or safety assistance;
  • change credentials, revoke permissions, or protect data;
  • request human review, correction, appeal, or compensation; or
  • understand a consequential decision made about them.

Downstream operators and customers

Notify downstream parties when they must disable a feature, change a configuration, update a model or dependency, preserve logs, repeat an assessment, or contact their own affected users.

Upstream providers

Notify model, data, infrastructure, evaluation, and tool providers when their component may be causal, exposed, or necessary to containment. Reports SHOULD include reproducible evidence while minimising unnecessary transfer of personal or confidential data.

Authorities and sector bodies

Notify competent regulators, data-protection authorities, market-surveillance authorities, sector regulators, law enforcement, emergency services, or national incident-response bodies when required or when their intervention is necessary to reduce harm.

Researchers and the public

Public disclosure is presumptively appropriate when:

  • the affected population cannot be reliably identified;
  • similar systems are likely to share the failure mode;
  • independent scrutiny is necessary to evaluate remediation;
  • users need information to assess prior decisions;
  • the incident affects a matter of substantial public interest; or
  • material claims previously made about the system are no longer supportable.

8. Minimum disclosure record

Every AIDB-1 through AIDB-3 incident MUST have a durable record. Fields may be marked unknown, under investigation, withheld, or not applicable, but MUST NOT be silently omitted.

Identification

  • stable incident identifier;
  • report version and publication timestamp;
  • disclosing organisation and accountable incident owner;
  • secure contact channel;
  • current severity and status;
  • jurisdictions and sectors implicated.

System context

  • system name and intended purpose;
  • developer, provider, deployer, and material integrators;
  • affected model, API, dataset, retrieval index, tool, policy, and interface versions;
  • deployment environment and relevant configuration;
  • dates of first deployment, last material change, and affected operating period;
  • reasonably foreseeable uses implicated by the incident.

Event chronology

  • earliest known occurrence;
  • detection time and detection method;
  • internal escalation time;
  • containment and notification milestones;
  • known latency between occurrence and detection;
  • time zone for every timestamp.

Impact

  • observed and reasonably foreseeable harms;
  • number and characteristics of affected people or entities;
  • known disproportionate effects on protected or vulnerable groups;
  • geographic and temporal scope;
  • reversibility and available remedy;
  • uncertainty bounds and assumptions used for estimates.

Technical and organisational analysis

  • triggering inputs, conditions, or system interactions;
  • expected versus observed behaviour;
  • causal status: confirmed, contributory, plausible, disputed, or unknown;
  • relevant training, evaluation, monitoring, oversight, and access-control failures;
  • whether the issue was reproducible and under what conditions;
  • indicators of malicious exploitation or coordinated misuse;
  • dependencies and downstream systems that may share exposure.

Response

  • containment measures completed or in progress;
  • model, data, software, policy, or workflow changes;
  • customer and user mitigations;
  • residual risk and known limitations of the fix;
  • validation performed before restoration;
  • rollback criteria and monitoring period;
  • remedy, appeal, or redress available to affected parties.

Disclosure controls

  • information withheld and the reason;
  • planned date for reconsidering withheld information;
  • legal or regulatory notices made;
  • coordination partners;
  • update schedule;
  • criteria for closure.

9. Evidence and reproducibility standard

AI incidents frequently involve stochastic outputs, changing models, inaccessible third-party components, and incomplete observability. A disclosure process must therefore preserve more than screenshots or selected transcripts.

For material incidents, the organisation SHOULD preserve where lawful and proportionate:

  • model and system version identifiers;
  • prompts, system instructions, retrieved context, tool calls, and outputs;
  • sampling parameters, seeds, and inference settings where available;
  • safety-policy and classifier versions;
  • relevant training, fine-tuning, and evaluation lineage;
  • timestamps, request identifiers, deployment region, and access pathway;
  • human interventions and override decisions;
  • monitoring alerts and audit logs;
  • known-good and known-bad comparison cases;
  • test harnesses, reproduction rate, and confidence intervals;
  • chain-of-custody records for evidence used in legal or regulatory proceedings.

Reports MUST state whether findings are:

  1. directly observed;
  2. reproduced under controlled conditions;
  3. inferred from logs or statistical evidence;
  4. reported by a third party but not independently verified; or
  5. hypothesised pending further investigation.

Where deterministic reproduction is impossible, the investigator SHOULD report the number of trials, model and system conditions, observed frequency, uncertainty, and factors known to alter the result.

10. Root-cause analysis

Root cause MUST NOT be reduced to "the model hallucinated" or "the user misused the system." Analysis SHOULD examine interacting causes across:

  • data: provenance, quality, representativeness, contamination, consent, and drift;
  • model: objective, architecture, training process, capability, alignment, and evaluation coverage;
  • system: retrieval, memory, tools, permissions, interfaces, and fallback behaviour;
  • operations: monitoring, change management, logging, access, and rollback readiness;
  • human factors: automation bias, workload, training, escalation design, and authority to intervene;
  • governance: ownership, incentives, risk acceptance, procurement, documentation, and assurance;
  • ecosystem: upstream dependencies, downstream adaptation, adversarial pressure, and cumulative impact.

The final report SHOULD identify both the initiating event and the control failures that allowed the event to produce harm.

11. Coordinated disclosure

AI incidents may expose vulnerabilities or capabilities that would be dangerous to publish immediately. Coordinated disclosure is justified only when temporary restriction of detail is necessary to enable mitigation.

An organisation using coordinated disclosure MUST:

  • acknowledge receipt through a published reporting channel;
  • establish a named coordinator and secure communication method;
  • provide the reporter with an initial assessment timeline;
  • preserve the reporter's contribution and preferred attribution;
  • avoid legal threats against good-faith research consistent with the published policy;
  • document the risk basis for delaying public detail;
  • share actionable mitigations with exposed parties as early as safely possible;
  • define a disclosure deadline and review any extension;
  • publish enough information for users to understand residual risk.

RFC 9116 can identify the reporting channel through security.txt, while ISO/IEC 29147 provides a mature model for receiving and coordinating vulnerability reports. Neither should be treated as sufficient by itself for non-security AI harms.

12. Disclosure quality tests

Before publication, the incident owner SHOULD be able to answer yes to each question:

  1. Can an affected person understand what happened and what to do next?
  2. Can a downstream operator determine whether its deployment is exposed?
  3. Can an independent investigator distinguish fact from hypothesis?
  4. Are the affected model and system versions identifiable?
  5. Does the chronology reveal detection and response latency?
  6. Are scale estimates accompanied by assumptions and uncertainty?
  7. Are disproportionate effects and vulnerable populations considered?
  8. Are mitigations specific, testable, and bounded by known limitations?
  9. Is withheld information identified with a defensible reason?
  10. Is there a date or trigger for the next update?

A notice that fails these tests may satisfy a communications objective while failing the purpose of incident disclosure.

13. Governance requirements

An organisation claiming AIDB conformance MUST:

  • assign executive accountability for AI incident management;
  • maintain a cross-functional response function spanning technical, security, privacy, safety, legal, communications, and domain expertise;
  • publish or contractually provide an intake channel for AI incident reports;
  • define severity authority and escalation paths before an incident occurs;
  • maintain an inventory linking AI systems to owners, versions, data, vendors, and downstream uses;
  • preserve incident evidence under documented retention and access rules;
  • test the process through exercises at least annually and after material architectural change;
  • track corrective actions to verified closure;
  • review recurring low-severity events for aggregate or systemic risk;
  • conduct a post-incident review that addresses control and governance failures, not only model behaviour.

The incident function MUST have authority to recommend suspension, rollback, capability restriction, user notification, and regulator engagement without waiting for normal product-release cycles.

14. Metrics that matter

Organisations SHOULD measure:

  • time from first occurrence to detection;
  • time from detection to accountable escalation;
  • time to containment;
  • time to first actionable external notice;
  • percentage of notices materially corrected after publication;
  • percentage of affected parties reached;
  • recurrence rate by failure mode;
  • corrective-action closure and validation rate;
  • incidents detected by users or external researchers rather than internal controls;
  • near misses aggregated into material risk findings;
  • disparities in impact and remedy across affected groups.

Raw incident counts MUST NOT be used as a standalone safety-performance measure. A rising count may reflect worsening controls, improved detection, increased deployment, or a healthier reporting culture.

15. Machine-readable reference record

Organisations SHOULD publish a machine-readable record alongside the human-readable notice. This minimal example is intentionally extensible:

{
  "schema": "https://reseni.io/schemas/aidb/v1",
  "incident_id": "AIDB-2026-0042",
  "version": "1.1",
  "status": "monitoring",
  "severity": "AIDB-2",
  "published_at": "2026-09-07T16:00:00Z",
  "updated_at": "2026-09-07T18:30:00Z",
  "organisation": {
    "name": "Example Organisation",
    "contact": "https://example.org/.well-known/security.txt"
  },
  "system": {
    "name": "Example Decision Support",
    "versions": ["model-2026-08-15", "orchestrator-4.2.1"],
    "role": "deployer"
  },
  "timeline": {
    "first_known_occurrence": "2026-09-03T09:15:00Z",
    "detected_at": "2026-09-06T14:40:00Z",
    "contained_at": "2026-09-07T11:20:00Z"
  },
  "impact": {
    "summary": "Incorrect prioritisation recommendations under defined input conditions.",
    "affected_count": {
      "estimate": 184,
      "lower_bound": 150,
      "upper_bound": 230
    },
    "reversibility": "partial"
  },
  "causal_status": "contributory",
  "mitigations": [
    "Disabled the affected recommendation path",
    "Initiated human review of decisions made during the exposure window"
  ],
  "residual_risk": "Historical decisions remain under review.",
  "next_update_due": "2026-09-14T16:00:00Z"
}

16. Adoption profile

An organisation can implement AIDB in four stages:

Stage 1 — Establish

Create accountable ownership, publish an intake channel, define severity levels, and connect AI incidents to existing security, privacy, safety, and business-continuity processes.

Stage 2 — Instrument

Add versioning, lineage, logging, evaluation records, downstream contact data, and evidence-preservation controls sufficient to investigate incidents.

Stage 3 — Exercise

Run scenario-based exercises involving stochastic failure, third-party model dependency, affected-user notification, regulatory reporting, and contested causality.

Stage 4 — Assure

Review incident decisions independently, test whether mitigations prevent recurrence, publish aggregate transparency data, and use findings to revise risk assessments and system design.

17. Relationship to established frameworks

AIDB is designed to complement, not replace:

These instruments address different objects and audiences. A management-system standard establishes organisational controls; a risk framework structures decision-making; a vulnerability standard coordinates remediation; and law creates jurisdiction-specific duties. AIDB supplies the operational bridge from an observed AI event to an evidence-based disclosure.

18. Research limitations and revision agenda

This baseline requires validation across sectors, jurisdictions, organisation sizes, and AI architectures. Priority research questions include:

  • which severity indicators best predict long-term and distributed harm;
  • how reporting thresholds should account for cumulative discrimination and low-frequency catastrophic risk;
  • how to quantify affected populations when model observability is limited;
  • when reproducibility thresholds become misleading for stochastic systems;
  • how to disclose dual-use capability incidents without increasing misuse;
  • how downstream deployers can report incidents when providers withhold model evidence;
  • how public incident repositories can preserve utility while protecting personal and security-sensitive information.

AIDB should evolve through documented versions informed by incident data, practitioner exercises, affected-community input, regulator guidance, and independent research.

Conclusion

AI incident disclosure should not begin with public relations and end with a patch note. It is an evidence discipline: identify the event, preserve the system state, classify harm, route actionable information, state uncertainty, correct the record, and verify that remediation works.

The minimum credible standard is straightforward: disclose early enough to reduce harm, precisely enough to support action, and honestly enough that an independent reader can understand both what is known and what remains unresolved.

AI Incident Disclosure Baseline (AIDB) v1.0 · Reseni Labs