Document
NIST AI 200-2 ipd, The TEVV-Athlon Framework for Evaluating AI Systems (Initial Public Draft)
Submitted
4 October 2026, in response to NIST's public comment request
Request items
2, 3, 4, 5 and 6
Sections
2.2, 2.3, 2.4, 3, 4.3.2 and 4.5
Comment type
Methodological clarification and high-consequence worked-example proposal

The issue

TEVV (test, evaluation, verification and validation) assesses whether an AI system meets its specifications and is sufficient for its intended use. For systems that can take consequential autonomous action, that finding is often used to support a further decision: whether the organisation permits the system to act. The comment asks NIST to keep that boundary explicit, especially as NIST names agentic systems among TEVV-Athlon's intended applications.

Three propositions should remain distinct:

Capability

The system can perform a function under tested conditions.

Trustworthiness evidence

The system exhibits measured properties such as validity, reliability, safety, robustness or security under defined conditions.

Operational permission

The organisation permits the system to perform a consequential action within a defined scope and subject to continuing conditions.

A TEVV result may inform the third, but should not be read as automatically establishing it. This matters most where evidence is conditional, environments change after evaluation, consequences are hard to reverse, or intervention must happen within a short decision window.

Flow diagram: TEVV evidence and its validity conditions inform an organizational decision, which sets a bounded operational scope (permitted actions, modes and assets, relevant conditions, duration or review horizon). The system operates while material changes are monitored. If the evidence is still decision-relevant, operation continues within the bounded scope; if not, targeted reassessment of the evidence or decision follows, and updated evidence can inform a revised decision.
Figure 1. Illustrative evidence-to-decision interface: TEVV evidence informs an organizational decision; the resulting bounded operational scope remains a separate organizational determination and may be revisited when evidence validity materially changes.

Central recommendation

TEVV should state not only what the evidence shows, but also the bounded operational scope the evidence can reasonably support and the conditions that keep that conclusion decision-relevant.

Where useful, that scope can be described by permitted actions, operational modes, affected assets, relevant system or environmental conditions, and a duration or review horizon. These fields are descriptive: they record what the evidence can support, not an enforcement or runtime permission mechanism. The proposal does not change TEVV-Athlon's four-stage structure or turn it into an authorization framework.

Five proposed changes

  1. Clarify the decision scope of TEVV evidence (Section 2.4). Reports should separate measured capability from the permission decision and state the assumptions, uncertainty and conditions that bound the evidence.
  2. Measure intervention, fallback, containment and withdrawal (Sections 2.2–2.3). Use existing Metrology Blocks for measures such as intervention feasibility, fallback effectiveness and permission-withdrawal latency where relevant.
  3. Test operational-mode transitions (Section 4.3.2). Realistic testing can cover moves between recommend-only, supervised, bounded-autonomous, degraded, suspended, fallback and restored modes.
  4. Link measurement validity to decision review (Section 4.5). Identify conditions under which a material change in validity would warrant targeted reassessment of the evidence or the decision it informed.
  5. Add a high-consequence worked example (Section 3). Complement the current low-impact chatbot example with a cyber-physical or critical-infrastructure case.

Illustrative example

The comment includes a vendor-neutral TEVV-Athlon table for a telecommunications function that detects network degradation and proposes or executes bounded traffic rerouting. It maps trustworthiness characteristics to Blocks, Events and Tools, including fault-detection quality, reroute quality, unsafe-action frequency, intervention feasibility, permission-withdrawal latency and fallback effectiveness. It also outlines a test design that deliberately invalidates the evidence supporting continued autonomous action while the system stays operational and connected.

The comment places the proposal alongside related NIST work, including NIST AI 800-4 on monitoring deployed AI and the NCCoE project on software and AI agent identity and authorization, and closes with three questions for NIST consideration. The full text, tables and references are in the PDF.

Download full comment (PDF) → Discuss this contribution

Disclosure (from the submission): No proprietary information. AI tools assisted with drafting and editorial revision. I am responsible for the submitted views and recommendations.

Citation: Arora, G. (2026). Extending TEVV-Athlon from Evaluation Evidence to Bounded Operational Permission: Considerations for Agentic, Autonomous, and High-Consequence AI Systems. Public comment on NIST AI 200-2 ipd, submitted 4 October 2026. Institute for Technology Stewardship.