# Methodology

## Research question

How often do public federal AI solicitation packages include safeguards needed to test, monitor,
govern, and eventually replace an AI system?

## Study frame

The frame contains public SAM.gov solicitations posted from January 1, 2024, through August 28,
2026. Included notice types are Solicitation, Combined Synopsis/Solicitation, Presolicitation,
and Sources Sought. The study measures language observable in the public package. It does not
claim that an agency lacks a practice internally when the language is not public.

Official FY2024–FY2026 archived Contract Opportunities exports were scanned locally. The current
frame contains 1,532 taxonomy-positive candidate notices: 570 from 2024, 555 from 2025, and 407
from 2026. Notice versions are retained separately rather than silently overwritten.

## AI procurement taxonomy

The candidate taxonomy covers artificial intelligence, machine learning, generative AI, large
language models, computer vision, natural-language processing, predictive modeling, intelligent
automation, and contextual `AI`/`AI-ML` expressions. Standalone `AI` requires linguistic context
to reject military part-code collisions. A 240-row weighted positive sample and 600-row weighted
negative sample—half drawn from deliberately oversampled near misses—support precision and recall
estimation.

At the project owner's direction, one versioned Codex semantic-review pass resolved every sampled
row as `yes` or `no`. The review used the complete archived candidate descriptions, explicit and
semantic AI tests, generic-technology-list exclusions, and acronym-collision checks. Final
estimates use stratum inclusion weights. Precision and recall intervals use stratified
finite-population variance; singleton strata receive the conservative binary variance bound of
0.25. This was not an independent human review, so no reviewer-agreement or kappa statistic is
reported.

## Documents and provenance

SAM’s public archived-resource service was queried for all 1,532 notices. It returned successfully
for 1,509 and returned 404 for 23. The resulting 4,699 public attachments were downloaded with no
unresolved download failures. Each source retains its URL, retrieval time, publisher, notice date,
stable identifier, version, original filename, MIME type, SHA-256 hash, extraction method, tool
version, confidence, and review state.

PDF, DOCX, XLSX/XLS, PPTX, RTF, HTML, JSON, plain text, legacy DOC, and bounded ZIP contents were
processed. All 353 scan-only PDFs completed OCR. One file has a retained extraction failure, and
seven are classified as non-text.

## Ten safeguards

Each notice package is evaluated for:

1. High-impact-use disclosure
2. Realistic testing and evaluation
3. Measurable performance requirements
4. Ongoing testing and monitoring
5. Government data and intellectual-property rights
6. Data/model portability and exit protection
7. Open APIs and interoperability
8. Licensing and lifecycle-pricing transparency
9. Model or feature change notification
10. Sunset, transition, and closeout provisions

Deterministic high-precision expressions nominate passages; they do not establish a reviewed score.
The validation frame contains 10,760 notice-category combinations across the 1,076 notices with
scoreable text. A stratified sample of 100 judgments per category oversamples rule-positive rows
and uses inverse-probability weights. A single Codex semantic-review pass inspected the extracted
text from each complete public package and resolved each row as `meets` or `does_not_meet`. The
review looked for the broader category concept together with binding solicitation language, not
only the deterministic retrieval expression or nominated excerpt.

Category estimates, rule precision/recall, and 95% confidence intervals are released only when all
1,000 sampled judgments are resolved. Reviewer agreement and Cohen's kappa are not calculated
because the release has one disclosed review pass. Singleton sampling strata use a conservative
variance bound rather than zero variance.

The headline GS Federal AI Procurement Safeguards Index is defined as 100 times the equal-weight
mean of the ten reviewed category prevalence estimates. It is a population-level summary, not a
claim that every notice has a reviewed notice-level score. Agency, notice-type, set-aside,
and policy-period tables use clearly labeled deterministic rule-detection rates; they are not
substituted for the Codex-reviewed overall category estimates.

The release-mode overall chart displays the same reviewed category estimates and 95% confidence
intervals used in the headline index. Pre/post, agency, and notice-level distribution charts are
explicitly labeled as deterministic rule-detection diagnostics.

## Comparisons

The descriptive tables compare exact agency names, notice types, set-aside status, and periods
before versus on/after April 3, 2025, the issue date of OMB M-25-22. Agency tables suppress groups
with fewer than ten candidate notices by default. The pre/post comparison is descriptive and is
not a causal estimate: the study does not assume immediate implementation, common agency mix, or
unchanged procurement composition.

Every table reports both the all-notice denominator and the scoreable-public-text denominator.
This prevents missing packages from being silently treated as affirmative or fully observed.

## Awards

Award linkage uses two exact-identifier paths:

- SAM Contract Awards API responses are accepted as candidates only when the returned
  `coreData.solicitationId` exactly matches the notice’s solicitation number.
- Official archived Award Notices are joined locally on normalized exact solicitation number,
  with agency and chronology checks.

Candidates are grouped to one opportunity-level decision while preserving every underlying PIID,
award notice, modification, amount, and query status. A single Codex review accepted or rejected
all 245 SAM API groups and 65 archive groups using exact identifier, agency, chronology, proposal,
and candidate-context checks. Multi-award vehicles,
modifications, ceilings, and obligations are not summed without a documented aggregation rule.
USAspending exact-PIID enrichment is implemented, but the live endpoint returned HTTP 503 in the
feasibility run; this does not change the solicitation-language index.

## Policy and source references

- [OMB M-25-22, issued April 3, 2025](https://www.whitehouse.gov/wp-content/uploads/2025/02/M-25-22-Driving-Efficient-Acquisition-of-Artificial-Intelligence-in-Government.pdf)
- [GAO-26-107859, released April 13, 2026](https://www.gao.gov/products/gao-26-107859)
- [SAM.gov Get Opportunities API](https://open.gsa.gov/api/get-opportunities-public-api/)
- [SAM.gov Contract Awards API](https://open.gsa.gov/api/contract-awards/)
- [SAM.gov archived Contract Opportunities data](https://sam.gov/data-services/Contract%20Opportunities/Archived%20Data)
- [USAspending API](https://api.usaspending.gov/docs/)

## Limitations

- Eligibility and safeguard judgments come from one automated semantic-review pass, not
  independent human replication; future human review could change individual decisions.
- Public attachments may omit contract clauses held elsewhere or restricted from public access.
- A rule non-match means “not found,” not “absent from agency practice.”
- Source documents vary greatly in structure and OCR quality.
- Category rules intentionally favor precision; the reviewed sample estimates false negatives.
- Agency comparisons are descriptive and can reflect procurement mix.
- Award links are contextual research outputs and are not needed to calculate the safeguards index.
