Artificial Intelligence Performance Evaluation Tool
Overview
What you need to judge this in 30 seconds
STTR · Phase: BOTH · Topic DAF26TZ06-NV006 · Solicitation 26.TZ
The proposed research effort aims to address critical gaps in the evaluation and validation of Artificial Intelligence (AI) systems within the Air Force by developing an Artificial Intelligence Performance Evaluation Tool. The tool will provide a standardized, adaptive, and robust framework capable of assessing the performance, reliability, and explainability of AI models, with a targeted application in sensor and information systems used in test and evaluation missions under the 412th Test Wing (412TW) of the Air Force Test Center (AFTC). This effort aligns with the Department of Defense's (DoD) thrust toward advancing Trusted AI and Autonomy and operationally focused Advanced Battle Management System (ABMS) initiatives. Problem/Unmet Need: Current AI evaluation processes lack standardization and reproducibility across varying operational environments, making it difficult for the AFTC to confidently assess AI models’ trustworthiness, reliability, and operational readiness. Furthermore, the escalation in complexity of AI-enabled systems has underscored the need for technologies that can quantitatively measure AI performance while adhering to DoD Responsible AI principles. These challenges prevent the Air Force from leveraging AI advancements at scale, impacting readiness and modernization goals. Opportunity/Desired Outcome: By creating a modular AI evaluation tool, this effort provides an opportunity to establish a scalable and adaptable solution that aligns with emerging DoD AI standards. The desired outcome is a measurable enhancement in the fidelity, speed, and consistency of AI performance assessments, reducing decision-making cycles and increasing mission readiness. The tool will enable practical implementation of AI validation in secure environments, ensuring compliance with applicable standards and operational conditions. Successful deployment will support the broader modernization priorities tied to mission-critical AI deployment and advanced autonomy. Approach The project will begin in Phase I with baseline research and feasibility studies, advancing into a scalable prototype in Phase II for integration and testing in operationally relevant environments. A summary of planned efforts by phase follows: Phase I: Feasibility and Concept Development (Initial TRL: 2; Target TRL: 4) Research: Conduct an initial landscape analysis of AI performance metrics, methodologies, and tools with a focus on gaps and challenges specific to AFTC’s mission space. Identify and assess algorithms for measuring real-time inference latency, adversarial resilience, and data drift specifically for different sensor models. This research must establish and utilize open-source or synthetically generated surrogate sensor datasets that are representative of AFTC flight test scenarios. Framework Design: Develop a software architecture for an adaptable AI performance evaluation framework designed to address reliability, robustness, and explainability metrics. Awardees will need to base their design on widely accepted DoD digital engineering and telemetry standards. The architecture must explicitly define the assumed data ingestion pipelines for surrogate sensor data and detail the computational methods that will be used to calculate reliability, robustness, and explainability metric. Simulation: Perform initial simulation tests to explore early-stage integration of explainability metrics and KPI identification methodologies. This will consist of running dry-runs within a localized software sandbox. The objective of these tests is to demonstrate that the proposed mathematical models can successfully process these inputs, identify Key Performance Indicators (KPIs) such as data drift, and output quantifiable explainability scores. Deliverables: A feasibility report describing how the tool can practically operate within the AFTC environment, a detailed framework design document, and initial simulation results demonstrating concept validity. This could inc…
- Category
- R&D
- Industry
- Needs verification
- Technology
- Needs verification
- Target stage
- Needs verification
- Project duration
- Needs verification
- Estimated preparation
- Needs verification
Funding
Award size and how it is paid
- Award range
- Needs verification
- Currency
- USD
- Total programme budget
- Needs verification
- Support type
- Needs verification
- Co-funding
- Needs verification
- Matching fund
- Needs verification
- Disbursement
- Needs verification
- Note
- -
Eligibility
Can we actually apply?
SBIR/STTR 은 미국 중소기업만 지원할 수 있습니다(Small Business Act 법정 요건). • 계열사를 포함해 상시 종업원 500명 이하 • 미국 시민 또는 영주권자 1인 이상이 50%를 초과해 직접 소유·지배 • 미국 내 사업장을 두고 주로 미국 내에서 사업을 영위할 것 • 수행책임자(PI)의 주된 근무처가 신청 기업일 것 출처: https://www.sbir.gov/faq/eligibility-requirements
- Company age
- Needs verification
- Employees
- 0명 ~ 500명
- Revenue limits
- Needs verification
- Consortium
- Needs verification
Location Requirements
Geography and legal-entity conditions
- Primary country
- 🇺🇸 United States
- Also eligible
- None
- Foreign companies
- No
- Local entity
- Required
- Location condition
- 미국 내 사업장을 두고 미국 시민·영주권자가 50%를 초과해 소유한 중소기업만 신청할 수 있습니다(법정 요건).
Required Documents
Required/optional · issuer · difficulty · validity · cautions
Document requirements were not captured (needs verification). Check the official announcement.
Application Process
How you apply
Application process not captured.
Evaluation
Review process and criteria
Evaluation process not captured.
Timeline
Announcement → intake → review → agreement → execution
Process steps were not captured.
Equity & Financial Terms
Equity, loans and contract conditions
- Equity required
- Needs verification
- Loan
- Non-dilutive grant
- Duplicate funding
- Needs verification
- IP ownership
- Needs verification
- Deliverable ownership
- Needs verification
- Exclusivity
- Needs verification
- Right of first refusal
- Needs verification
- Audit & settlement
- Needs verification
Risk Intelligence
Should we apply at all? — five dimensions plus an overall score (lower is safer)
No risk assessment yet. Re-analyse from Admin.
IP / Idea Protection
How much of your technology and idea you must disclose
No IP assessment yet.
AI Analysis
Pros · cons · difficulty · competition · attractiveness
From paperwork and review stages
Heuristic from award size and type
Amount, terms and IP combined
Weighted average of five dimensions
- Needs verification
- Needs verification
- Nothing flagged
Similar Programs
Comparable by type, country and technology
- Counter Adversial GPS Jamming
🇺🇸 United States · R&D program · U.S. Department of War — USAF
Needs verificationD-14 - Ground and Air Launched Drone Swarms Create a Self-Protecting Perimeter Using Autonomous AI
🇺🇸 United States · R&D program · U.S. Department of War — USAF
Needs verificationD-14 - Space Domain Operational Environment Assessment
🇺🇸 United States · R&D program · U.S. Department of War — USAF
Needs verificationD-14 - Low Noise & Low-SWaP High-Repetition Rate Mode-Locked Lasers for High-Speed Photonics
🇺🇸 United States · R&D program · U.S. Department of War — USAF
Needs verificationD-14 - Low-Cost Multi-Function Ku-Band RF System for Prompt Strike Weapons
🇺🇸 United States · R&D program · U.S. Department of War — USAF
Needs verificationD-14
Source & Provenance
Every value carries its source URL and fetch time
- Official URL
- https://www.dodsbirsttr.mil/topics-app/
- Application URL
- https://www.dodsbirsttr.mil/topics-app/
- Last fetched
- 2026.10.07
- Human review
- Not reviewed — check the official announcement
- First seen
- 2026.10.07
- Data status
- Auto-collected, not reviewed
- AI enrichment
- Needs verification
- PRIMARY
- HTMLSBIR/STTR 자격요건 (법정)
Fetched 2h ago
- HTML발주 기관 공식 공고
Fetched 2h ago
- · 2h ago — 최초 수집 (e138e647)