Original Research | | 15 min read
We Reviewed 1,532 Federal AI Solicitations. The Safeguards Score Was 31.5 Out of 100.
Key Takeaways
Federal AI buying is moving faster than its public safeguards
31.5 out of 100
The equal weight mean of ten weighted reviewed safeguard estimates.
1,532 public notices
1,076 notices had scoreable public package text. The remainder limits what can be observed.
59.7% performance metrics
Measurable performance language was the strongest reviewed category estimate.
3.3% high impact disclosure
Public language addressing high impact AI determinations was the weakest reviewed category estimate.
Federal AI buying is moving faster than the public safeguards wrapped around it.
GS Consulting conducted and reviewed this original research. We examined 1,532 public federal solicitation notices in a fixed study frame from January 1, 2024 through August 28, 2026. The resulting Federal AI Procurement Safeguards Index scored 31.5 out of 100.
The problem is not that federal buyers are ignoring AI risk. The problem is that the public acquisition record often does not show enough about how that risk will be tested, measured, monitored, priced, changed, transferred, and closed out.
If 31.5 sounds abstract, think of a report card with ten equally important subjects. We estimated how often each safeguard appeared in the public packages we could score, then averaged the ten reviewed estimates. The score describes public evidence. It is not an agency grade, a legal compliance score, or a prediction that a project will fail.
This research belongs in the larger GovCon AI evidence library. Teams turning the findings into acquisition and operating controls should also use GS Consulting's AI Governance, Risk, and Oversight guidance.
Explore the evidence behind the headline.
Filter all 1,532 notices in the research explorer or download the full report, data, figures, methods, and reporter package.
Federal AI Procurement Safeguards: The Short Answer
Public solicitations regularly described what an AI system should do. They were less consistent about the protections that make the purchase durable and governable after award.
Performance metrics were the most visible reviewed safeguard, with an estimated prevalence of 59.7 percent. High impact disclosure was the least visible at 3.3 percent. Feature change notification was next at 7.2 percent. That gap matters because an AI service can change after evaluation, and a buyer cannot govern a material change it is never told about.
The practical lesson is simple: a statement of work should not stop at features. It should define the evidence the buyer will receive before acceptance, during operation, after a material change, and at contract exit.
What the 31.5 Out of 100 Index Shows
The reviewed headline index is an equal weight average of ten safeguard estimates. Each category gets the same weight. That prevents a frequently mentioned category, such as performance metrics, from hiding a rarely mentioned category, such as high impact disclosure.
The reviewed estimates correct for errors found when we validated the retrieval rules. A deterministic rule is a repeatable text search: the same rule applied to the same text produces the same match. Confidence intervals show a reasonable uncertainty range around each reviewed estimate. The index is therefore different from the raw rule count shown in the explorer and several diagnostic charts.
Do not subtract 31.5 from 100 and call the result a failure rate. Public package text is incomplete. Of the 1,532 notices in the candidate frame, 456 did not have enough retrieved public text to score. Some safeguards may also exist in internal plans, nonpublic attachments, evaluation records, or later contract documents.
The Ten Safeguards, in Plain English
- Performance metrics — 59.7%. Does the package say how useful, accurate, timely, reliable, or otherwise successful performance will be measured?
- Data and IP rights — 43.3%. Does it explain who can use the data, models, outputs, configurations, and other intellectual property?
- Portability and exit — 41.9%. Can the government retrieve its data and move to another solution without being trapped?
- Pricing transparency — 38.8%. Can the buyer see the units, assumptions, usage drivers, and likely cost changes?
- Test and evaluation — 38.8%. Does the package define how the AI will be tested before the buyer accepts it?
- Interoperability — 32.9%. Can the system exchange data and work with required government tools, formats, and interfaces?
- Ongoing monitoring — 25.9%. Does the package require operating signals, drift checks, incidents, and review after deployment?
- Sunset and closeout — 23.5%. Does it define what happens to data, access, records, models, and services when the work ends?
- Feature change notification — 7.2%. Must the vendor tell the buyer before a material model, feature, data, or service change?
- High impact disclosure — 3.3%. Does the package address whether the AI use is high impact and what additional controls follow?
These labels turn a broad phrase such as “responsible AI” into questions a contracting officer, program manager, evaluator, security lead, lawyer, and vendor can actually answer.
Before and After OMB Memorandum M-25-22
OMB Memorandum M-25-22, issued April 3, 2025, addresses efficient AI acquisition and includes themes such as testing, transparency, portability, avoiding vendor lock in, and ongoing monitoring.
We compared raw deterministic rule match rates before the memo with rates on or after its issue date. Ongoing monitoring rose 5.6 percentage points and interoperability rose 5.7 points. Performance metrics fell 7.1 points. Other categories moved less.
The memo did not necessarily cause these movements. The mix of agencies, notice types, package availability, timing, and missing text changed across the periods. Treat this chart as a question generator, not a policy scorecard.
The Agency View Is a Diagnostic, Not a Ranking
Among agencies with at least ten candidate notices, NASA had the highest average raw deterministic index at 16.7, followed by the General Services Administration at 15.9 and Veterans Affairs at 15.1. Those numbers are not the reviewed headline index.
Do not read this as a league table. Defense accounts for 770 candidate notices, while Agriculture accounts for 13. The acquisitions are different, the public packages are different, and the amount of retrievable text is different.
A separate 2026 GAO review of 13 AI acquisitions at Defense, Homeland Security, GSA, and Veterans Affairs reported that agencies were not systematically collecting lessons from those acquisitions. Our research asks a different question, but the findings point in the same practical direction: buyers need reusable acquisition evidence, not one-time language that disappears inside a procurement file.
What the Public Package Could and Could Not Show
The raw deterministic score counts the number of safeguard categories with a rule match and multiplies that count by ten. It helps test retrieval. It should not replace the weighted reviewed index.
There were 880 notices in the zero match group. That number needs context. More than half of that group, 456 notices, lacked scoreable public text. “No match” is not the same as “no safeguard.”
The study started with 1,532 unique candidate notices. It completed 1,509 successful notice resource lookups and recorded 4,699 attachments. A total of 1,076 notices had scoreable public package text. Of those, 652 had at least one deterministic safeguard match.
What Federal Buyers Should Require
A buyer does not need ten more policy paragraphs. A buyer needs ten evidence requirements attached to the acquisition, evaluation, acceptance, and operating plan.
- Make the impact decision explicit. State whether the use could materially affect rights, safety, benefits, employment, access, or another high consequence outcome. Name the decision owner and the extra review that follows.
- Define realistic tests. Use representative data, difficult cases, failure cases, prohibited actions, security tests, and repeat runs. Set pass, fail, and review thresholds before evaluation.
- Turn performance into acceptance evidence. Name each metric, population, data source, threshold, reporting period, exception, and remedy.
- Monitor after deployment. Require drift, error, incident, override, latency, availability, and cost signals that match the actual use.
- Set data and IP boundaries. State who can use government data, prompts, outputs, feedback, configurations, model improvements, and derived artifacts.
- Design the exit before award. Require export formats, transition help, deletion evidence, credential removal, record retention, and a tested handoff path.
- Specify interoperability. Name required interfaces, standards, identity patterns, logs, data formats, and government systems.
- Expose the cost model. Separate fixed, usage, integration, support, model, storage, data transfer, and transition charges. Define how unit prices can change.
- Control material changes. Require notice before a model, feature, training source, hosting service, subprocessor, data practice, or safety control changes.
- Close the contract cleanly. Define what stops, what transfers, what remains accessible, what is destroyed, and who certifies completion.
Features are easy to buy. Evidence is harder.
GS Consulting helps federal and regulated teams translate AI risk into solicitation language, evaluation evidence, acceptance criteria, monitoring, and exit controls.
Review Your AI Procurement EvidenceWhat Contractors Should Prepare Before the RFP Asks
The same ten safeguards create a useful proposal evidence plan. A contractor that waits for a perfectly worded requirement will struggle to assemble credible proof during a short response window.
- Maintain a plain-language description of the AI use, impact boundary, prohibited uses, and accountable owners.
- Keep versioned test results tied to the offered model, data, configuration, tools, and operating environment.
- Show metric definitions, current baselines, failure thresholds, and how reported performance can be audited.
- Prepare monitoring samples, incident workflow, change notices, release notes, and customer notification timelines.
- Map data rights, model rights, output rights, training restrictions, subprocessors, and retention practices.
- Demonstrate export, migration, transition support, access removal, deletion, and closeout evidence.
- Explain unit economics in terms a program and contracting team can test against expected use.
The strongest response will not promise that the AI is safe, accurate, or compliant in every setting. It will show the buyer exactly what was tested, what remains uncertain, what will be monitored, and what happens when the system changes.
How GS Consulting Conducted and Reviewed the Research
GS Consulting built a fixed candidate frame from public SAM.gov notices dated January 1, 2024 through August 28, 2026. The release contains 1,532 unique notices. We classified the notices, resolved public notice resources, collected attachment records, extracted available public package text, applied ten documented safeguard retrieval rules, and reviewed the results.
The notice classifier validation produced 95.6 percent precision and 96.3 percent recall. In plain English, precision asks how often the classifier correctly identified a relevant notice. Recall asks how often it found the relevant notices available in the validation set.
The headline category estimates use reviewed validation results and weighting. The overall 31.5 index is the equal weight mean of the ten reviewed estimates. Confidence intervals quantify sampling uncertainty. The raw deterministic counts in the explorer and diagnostic charts answer a narrower retrieval question.
Review disclosure: GS Consulting conducted and reviewed this research. The labels were produced through one documented, versioned Codex semantic review completed at the project owner's direction. This was not independent human replication. The release does not claim inter-rater agreement or report a kappa statistic.
Award review materials are preserved in the complete release package, but award values are not aggregated. The source fields were not reliable enough to support a defensible total. That restraint is part of the method: missing or inconsistent data should not become a polished number merely because the number would be interesting.
For the full sampling logic, category rules, weighting, confidence intervals, hashes, validation records, and limitations, read the methodology and data dictionary.
Research Explorer and Downloads
Reporters, researchers, buyers, vendors, and reviewers can inspect the evidence at the level they need.
- Interactive research explorer. Filter all 1,532 candidate notices by agency, notice type, policy period, search term, and safeguard match.
- PDF report. Download the seven-page release report.
- Executive summary and reporter brief. Use the concise findings, definitions, and citation guidance.
- Methodology and data dictionary. Review the study frame, labels, rules, weighting, and fields.
- Notice-level scores CSV and evidence CSV. Reproduce notice views and inspect category evidence.
- Reviewed overall estimates CSV. Use the published category estimates and confidence intervals.
- Branded figure bundle. Download SVG and PNG versions of all five website figures.
- Complete research package. Download the full publication release, review tools, manifests, data, charts, explorer, and PDF.
- Corrections log and version history. Check changes before citing or reusing the data.
Recommended citation: GS Consulting, GS Federal AI Procurement Safeguards Index, Version 1.0.0, August 29, 2026.
Sources, Scope, and Caveats
- Office of Management and Budget Memorandum M-25-22, issued April 3, 2025.
- GAO-26-107859, Artificial Intelligence: Selected Agencies Need to Improve Acquisition Processes, released April 13, 2026.
- GSA SAM.gov Get Opportunities Public API.
- SAM.gov archived Contract Opportunities data.
- GSA SAM.gov Contract Awards API.
- USAspending API documentation.
This research measures observable public solicitation packages. It does not measure every internal agency document, evaluation, control, contract modification, or operating practice. A nonmatch does not prove absence.
The before and after policy comparison is descriptive, not causal. The agency chart is a retrieval diagnostic, not a performance ranking. The index is GS Consulting original research, not an official OMB, GAO, GSA, agency, legal, regulatory, or procurement determination.
Frequently Asked Questions
What is the Federal AI Procurement Safeguards Index?
It is a GS Consulting original research index that measures how often ten practical AI safeguards appeared in observable public federal solicitation packages. The reviewed headline index is the equal weight average of ten weighted category estimates.
What does a score of 31.5 out of 100 mean?
The ten reviewed safeguard estimates averaged 31.5 percent when each category received equal weight. It is not an agency grade, a legal compliance score, a project failure rate, or proof that 68.5 percent of safeguards were absent from internal acquisition work.
How many federal AI solicitation notices did GS Consulting review?
The fixed candidate frame contains 1,532 unique public notices from January 1, 2024 through August 28, 2026. Of those, 1,076 had enough retrieved public package text to score.
Does a low public evidence score mean an agency had no safeguard?
No. The research measures observable public language. A safeguard may have existed in internal documents, nonpublic attachments, later contract actions, or agency practice. A rule nonmatch also does not prove absence.
Was the research independently human reviewed?
No. GS Consulting conducted and reviewed the research using one documented, versioned Codex semantic review completed at the project owner's direction. The release does not claim independent human replication, inter-rater agreement, or a kappa statistic.
Can reporters cite and download the research?
Yes. Use the report, reporter brief, methods, data, figures, corrections log, version history, explorer, and complete package linked above. Cite Version 1.0.0 dated August 29, 2026.
The Bottom Line
The problem is not a lack of AI ambition. The problem is a thin public evidence layer around what gets tested, measured, monitored, changed, transferred, priced, and closed out.
Buyers can fix that by writing evidence into the acquisition. Contractors can prepare by proving those safeguards before the solicitation forces the issue. Reporters and researchers can use the downloadable record to test, challenge, and extend these findings.
Continue Reading
- AI Procurement Regulations for Government Contractors
- AI Disclosure in Federal Contracts
- GovCon AI Insights Hub
- AI Governance, Risk, and Oversight
Turn the research into better acquisition evidence.
GS Consulting helps teams design AI procurement, evaluation, governance, monitoring, and transition evidence that can survive review and operate after award.
Start the Conversation