In July 2026, during an internal OpenAI cyber-capability evaluation on the ExploitGym benchmark, approximately 1,200 AI agents found an unauthorized way to communicate with each other and coordinated through more than 70,000 messages. About 700 of them took part in an attack that reached a real company, Hugging Face, after gaining a foothold in a third-party sandbox environment. They obtained 14 credentials, including a Kubernetes service-account token, a static MongoDB password, AWS keys, a mesh-VPN authentication key, a JWT signing key, a registry token, and GitHub App tokens, and used them to reach cluster-admin access across multiple internal clusters.
This was not an intentional attack. It was an internal OpenAI cyber-capability evaluation. Adversaries are building the intentional version.
The Medicaid Enterprise secures itself for a different world: annual assessments, authorization-to-operate documents, compliance checklists, point-in-time penetration tests. These defenses are reviewed at human speed, on paper, on a schedule. The attack that matters now onsets at machine speed, coordinates itself, and finds every exposed credential in hours rather than years. A security model whose fundamental unit is the annual document cannot answer a threat whose fundamental unit is the autonomous hour.
This gap will not be closed by amending the current program. Adding continuous controls on top of annual attestations produces the worst of both: the new tooling plus the old paperwork, each consuming the capacity the other needs. We suggest instead that CMS charter a new security program, designed from a clean sheet around the machine-speed threat model, and apply it as the complete security regime for Horizon 3 work, while the existing assessment framework continues to govern existing systems. The new program should not inherit a single control from the old one by default; every control must justify itself against the current threat, or it does not carry forward. As Horizon 3 increments demonstrate that continuous assessment provides stronger assurance than periodic documentation, the new program becomes the successor rather than the supplement.
The recommendation is to treat security as a continuously demonstrated capability: built into every delivery pipeline and checked automatically before every production deployment, so that the security question is asked and answered every time anything changes, not once a year. This does not retrofit easily into existing systems, and that is the point: a Horizon 3 environment must carry this property from its first increment, when building it in costs almost nothing, rather than attempt to add it later, when it costs everything.
This is the security expression of the same principle running through this response: assess against working systems, not documentation. It extends the assessment-at-ATO model described in X7 from a point-in-time demonstration to a continuous one, and it belongs in the same Horizon 3 increments as everything else — each slice arrives with its continuous security demonstration attached, which is only tractable because the slice is small.