Original Research | | 15 min read

Federal AI Procurement Safeguards: What 1,532 Public Notices Reveal


GS Consulting Federal AI Procurement Safeguards Index for 1,532 public notices, including 1,076 with scoreable public text
GS Consulting original research

Key Takeaways

Federal AI buying is moving faster than its public safeguards

Headline index

31.5 out of 100

The equal weight mean of ten weighted reviewed safeguard estimates.

Study frame

1,532 public notices

1,076 notices had scoreable public package text. The remainder limits what can be observed.

Most visible

59.7% performance metrics

Measurable performance language was the strongest reviewed category estimate.

Least visible

3.3% high impact disclosure

Public language addressing high impact AI determinations was the weakest reviewed category estimate.

Federal AI buying is moving faster than the public safeguards wrapped around it.

We analyzed 1,532 public notices; 1,076 contained scoreable public text. The fixed study frame covers public federal solicitation notices dated January 1, 2024 through August 28, 2026. The sampled semantic review produced a weighted Federal AI Procurement Safeguards Index of 31.5 out of 100.

The problem is not that federal buyers are ignoring AI risk. The problem is that the public acquisition record often does not show enough about how that risk will be tested, measured, monitored, priced, changed, transferred, and closed out.

If 31.5 sounds abstract, think of a report card with ten equally important subjects. We estimated how often each safeguard appeared in the public packages we could score, then averaged the ten reviewed estimates. The score describes public evidence. It is not an agency grade, a legal compliance score, or a prediction that a project will fail.

This research belongs in the larger GovCon AI evidence library. Teams turning the findings into acquisition and operating controls should also use GS Consulting's AI Governance, Risk, and Oversight guidance.

Preview of the interactive Federal AI Procurement Safeguards Explorer with filters and research statistics Launch the interactive explorer Interactive research tool Filter the evidence behind the 31.5 score. Search the 1,532-notice frame. Narrow the results by agency, notice type, policy period, and safeguard rule match. Then inspect the public-text diagnostic behind the numbers. Open the Explorer
Use the explorer for the notice-level rule-match diagnostic. The weighted findings and methodology summary appear on this page.
GS Consulting research snapshot showing 1,532 public notices analyzed, 1,076 notices with scoreable public text, a 31.5 out of 100 safeguards index, and a wide gap between performance metrics and high impact disclosure evidence
Research snapshot. GS Consulting analyzed the fixed notice frame and reviewed documented semantic samples. Open the infographic for a full size view.

Federal AI Procurement Safeguards: The Short Answer

Public solicitations regularly described what an AI system should do. They were less consistent about the protections that make the purchase durable and governable after award.

Performance metrics were the most visible reviewed safeguard, with an estimated prevalence of 59.7 percent. High impact disclosure was the least visible at 3.3 percent. Feature change notification was next at 7.2 percent. That gap matters because an AI service can change after evaluation, and a buyer cannot govern a material change it is never told about.

The practical lesson is simple: a statement of work should not stop at features. It should define the evidence the buyer will receive before acceptance, during operation, after a material change, and at contract exit.

What the 31.5 Out of 100 Index Shows

The reviewed headline index is an equal weight average of ten safeguard estimates. Each category gets the same weight. That prevents a frequently mentioned category, such as performance metrics, from hiding a rarely mentioned category, such as high impact disclosure.

The weighted estimates adjust for errors observed in sampled semantic review of the retrieval rules. A deterministic rule is a repeatable text search: the same rule applied to the same text produces the same match. Confidence intervals show a reasonable uncertainty range around each estimate. The index is therefore different from the rule-match diagnostic shown in the explorer and several diagnostic charts.

Weighted reviewed estimates for ten federal AI procurement safeguards with 95 percent confidence intervals
Figure 1. Weighted reviewed estimates across ten safeguard categories. Open the figure for a full size version. Whiskers show 95 percent confidence intervals.

Do not subtract 31.5 from 100 and call the result a failure rate. Public package text is incomplete. Of the 1,532 notices in the candidate frame, 456 did not have enough retrieved public text to score. Some safeguards may also exist in internal plans, nonpublic attachments, evaluation records, or later contract documents.

The Ten Safeguards, in Plain English

  • Performance metrics — 59.7%. Does the package say how useful, accurate, timely, reliable, or otherwise successful performance will be measured?
  • Data and IP rights — 43.3%. Does it explain who can use the data, models, outputs, configurations, and other intellectual property?
  • Portability and exit — 41.9%. Can the government retrieve its data and move to another solution without being trapped?
  • Pricing transparency — 38.8%. Can the buyer see the units, assumptions, usage drivers, and likely cost changes?
  • Test and evaluation — 38.8%. Does the package define how the AI will be tested before the buyer accepts it?
  • Interoperability — 32.9%. Can the system exchange data and work with required government tools, formats, and interfaces?
  • Ongoing monitoring — 25.9%. Does the package require operating signals, drift checks, incidents, and review after deployment?
  • Sunset and closeout — 23.5%. Does it define what happens to data, access, records, models, and services when the work ends?
  • Feature change notification — 7.2%. Must the vendor tell the buyer before a material model, feature, data, or service change?
  • High impact disclosure — 3.3%. Does the package address whether the AI use is high impact and what additional controls follow?

These labels turn a broad phrase such as “responsible AI” into questions a contracting officer, program manager, evaluator, security lead, lawyer, and vendor can actually answer.

Before and After OMB Memorandum M-25-22

OMB Memorandum M-25-22, issued April 3, 2025, addresses efficient AI acquisition and includes themes such as testing, transparency, portability, avoiding vendor lock in, and ongoing monitoring.

We compared raw deterministic rule match rates before the memo with rates on or after its issue date. Ongoing monitoring rose 5.6 percentage points and interoperability rose 5.7 points. Performance metrics fell 7.1 points. Other categories moved less.

Raw deterministic safeguard rule match rates before and on or after OMB Memorandum M-25-22
Figure 2. Raw deterministic rule match rates before and on or after April 3, 2025. The comparison is descriptive, not causal.

The memo did not necessarily cause these movements. The mix of agencies, notice types, package availability, timing, and missing text changed across the periods. Treat this chart as a question generator, not a policy scorecard.

The Agency View Is a Diagnostic, Not a Ranking

Among agencies with at least ten candidate notices, NASA had the highest average rule-match diagnostic at 16.7, followed by the General Services Administration at 15.9 and Veterans Affairs at 15.1. Those numbers are not the weighted headline index, and they do not measure whether an agency had safeguards.

Average rule-match diagnostic by agency for agencies with at least ten candidate notices
Figure 3. Average rule-match diagnostic for agencies with at least ten candidate notices. A zero or low value means qualifying language was not evident in the retrieved public text; it does not mean the agency had no safeguards.

Do not read this as a league table. Defense accounts for 770 candidate notices, while Agriculture accounts for 13. The acquisitions are different, the public packages are different, and the amount of retrievable text is different.

A separate 2026 GAO review of 13 AI acquisitions at Defense, Homeland Security, GSA, and Veterans Affairs reported that agencies were not systematically collecting lessons from those acquisitions. Our research asks a different question, but the findings point in the same practical direction: buyers need reusable acquisition evidence, not one-time language that disappears inside a procurement file.

What the Public Package Could and Could Not Show

The rule-match diagnostic counts the number of safeguard categories with a deterministic match and multiplies that count by ten. It helps test retrieval. It should not replace the weighted index or be treated as an agency or solicitation safeguards score.

Distribution of 1,532 notices by the number of deterministic safeguard rule matches
Figure 4. Rule-match diagnostic distribution. The zero group includes 456 notices without scoreable public text.

There were 880 notices in the zero match group. That number needs context. More than half of that group, 456 notices, lacked scoreable public text. “No match” is not the same as “no safeguard.”

Coverage funnel from 1,532 candidate notices to 652 with at least one deterministic safeguard match
Figure 5. Public package coverage from candidate notice through scoreable text and deterministic matches.

The study started with 1,532 unique candidate notices. It completed 1,509 successful notice resource lookups and recorded 4,699 attachments. A total of 1,076 notices had scoreable public package text. Of those, 652 had at least one deterministic safeguard match.

What Federal Buyers Should Require

A buyer does not need ten more policy paragraphs. A buyer needs ten evidence requirements attached to the acquisition, evaluation, acceptance, and operating plan.

  1. Make the impact decision explicit. State whether the use could materially affect rights, safety, benefits, employment, access, or another high consequence outcome. Name the decision owner and the extra review that follows.
  2. Define realistic tests. Use representative data, difficult cases, failure cases, prohibited actions, security tests, and repeat runs. Set pass, fail, and review thresholds before evaluation.
  3. Turn performance into acceptance evidence. Name each metric, population, data source, threshold, reporting period, exception, and remedy.
  4. Monitor after deployment. Require drift, error, incident, override, latency, availability, and cost signals that match the actual use.
  5. Set data and IP boundaries. State who can use government data, prompts, outputs, feedback, configurations, model improvements, and derived artifacts.
  6. Design the exit before award. Require export formats, transition help, deletion evidence, credential removal, record retention, and a tested handoff path.
  7. Specify interoperability. Name required interfaces, standards, identity patterns, logs, data formats, and government systems.
  8. Expose the cost model. Separate fixed, usage, integration, support, model, storage, data transfer, and transition charges. Define how unit prices can change.
  9. Control material changes. Require notice before a model, feature, training source, hosting service, subprocessor, data practice, or safety control changes.
  10. Close the contract cleanly. Define what stops, what transfers, what remains accessible, what is destroyed, and who certifies completion.

Features are easy to buy. Evidence is harder.

GS Consulting helps federal and regulated teams translate AI risk into solicitation language, evaluation evidence, acceptance criteria, monitoring, and exit controls.

Review Your AI Procurement Evidence

What Contractors Should Prepare Before the RFP Asks

The same ten safeguards create a useful proposal evidence plan. A contractor that waits for a perfectly worded requirement will struggle to assemble credible proof during a short response window.

  • Maintain a plain-language description of the AI use, impact boundary, prohibited uses, and accountable owners.
  • Keep versioned test results tied to the offered model, data, configuration, tools, and operating environment.
  • Show metric definitions, current baselines, failure thresholds, and how reported performance can be audited.
  • Prepare monitoring samples, incident workflow, change notices, release notes, and customer notification timelines.
  • Map data rights, model rights, output rights, training restrictions, subprocessors, and retention practices.
  • Demonstrate export, migration, transition support, access removal, deletion, and closeout evidence.
  • Explain unit economics in terms a program and contracting team can test against expected use.

The strongest response will not promise that the AI is safe, accurate, or compliant in every setting. It will show the buyer exactly what was tested, what remains uncertain, what will be monitored, and what happens when the system changes.

How GS Consulting Conducted and Reviewed the Research

GS Consulting built a fixed candidate frame from public SAM.gov notices dated January 1, 2024 through August 28, 2026. The frame contains 1,532 unique notices, and 1,076 contained enough retrieved public package text to score. We classified the notices, resolved public notice resources, collected attachment records, extracted available public package text, and applied ten documented safeguard retrieval rules.

The notice classifier validation produced 95.6 percent precision and 96.3 percent recall. In plain English, precision asks how often the classifier correctly identified a relevant notice. Recall asks how often it found the relevant notices available in the validation set.

The headline category estimates use sampled semantic review results and weighting. The overall 31.5 index is the equal weight mean of the ten weighted estimates. Confidence intervals quantify sampling uncertainty. The deterministic counts in the explorer and diagnostic charts answer a narrower retrieval question.

Review disclosure: The semantic review evaluated documented samples rather than every notice. GS Consulting conducted that review; a separate review team did not independently replicate it. The release does not claim inter-rater agreement or report a kappa statistic.

Award values are not aggregated because the available source fields were not reliable enough to support a defensible total. That restraint is part of the method: missing or inconsistent data should not become a polished number merely because the number would be interesting.

Research Explorer

The interactive research explorer provides a read-only view of the 1,532-notice frame. Filter by agency, notice type, policy period, search term, and safeguard rule match.

The explorer shows a retrieval diagnostic, not an agency grade or a finding that a named solicitation lacked safeguards. Downloadable research files are not published.

Recommended citation: GS Consulting, GS Federal AI Procurement Safeguards Index, Version 1.0.0, August 29, 2026.

Sources, Scope, and Caveats

This research measures observable public solicitation packages. It does not measure every internal agency document, evaluation, control, contract modification, or operating practice. A nonmatch means a safeguard was not evident in the retrieved public acquisition text under the documented rule; it does not prove absence.

The before and after policy comparison is descriptive, not causal. The agency chart is a rule-match diagnostic, not a performance ranking. A notice-level or agency-level zero does not mean the named agency or solicitation had no safeguards. The index is GS Consulting original research, not an official OMB, GAO, GSA, agency, legal, regulatory, or procurement determination.

Frequently Asked Questions

What is the Federal AI Procurement Safeguards Index?

It is a GS Consulting original research index that measures how often ten practical AI safeguards appeared in observable public federal solicitation packages. The reviewed headline index is the equal weight average of ten weighted category estimates.

What does a score of 31.5 out of 100 mean?

The ten weighted safeguard estimates averaged 31.5 percent when each category received equal weight. It is not an agency grade, a legal compliance score, or a project failure rate. The remaining share does not prove safeguards were absent from internal acquisition work; it means they were not evident in the reviewed public acquisition text under this method.

How many public notices did GS Consulting analyze?

The fixed candidate frame contains 1,532 unique public notices from January 1, 2024 through August 28, 2026. Of those, 1,076 contained enough retrieved public package text to score. The semantic review used documented samples rather than every notice.

Does a low public evidence score mean an agency had no safeguard?

No. The research measures observable public language. A safeguard may have existed in internal documents, nonpublic attachments, later contract actions, or agency practice. A rule nonmatch also does not prove absence.

Was the research independently human reviewed?

No. GS Consulting analyzed the fixed notice frame and conducted a sampled semantic review. A separate review team did not independently replicate the work. The release does not claim inter-rater agreement or report a kappa statistic.

Can reporters cite the research?

Yes. Cite GS Consulting, GS Federal AI Procurement Safeguards Index, Version 1.0.0, August 29, 2026. The public explorer provides a read-only view of the rule-match diagnostic; downloadable research files are not published.

The Bottom Line

The problem is not a lack of AI ambition. The problem is a thin public evidence layer around what gets tested, measured, monitored, changed, transferred, priced, and closed out.

Buyers can fix that by writing evidence into the acquisition. Contractors can prepare by proving those safeguards before the solicitation forces the issue. Reporters and researchers can use the public explorer and cited source systems to test and challenge these findings.

Continue Reading

Turn the research into better acquisition evidence.

GS Consulting helps teams design AI procurement, evaluation, governance, monitoring, and transition evidence that can survive review and operate after award.

Start the Conversation

© GS Consulting, LLC . All Rights Reserved | For more information, contact us at info@gsconsultingllc.com. Image credit: ©iStock.com/Vertigo3d. Privacy Policy | Terms of Use