incidents / hf-2026-07

hugging face production intrusion by the artifactory swarm

confirmed breach final report
period2026-07-11 to 2026-07-14
discovered byvictim
categoriessandbox boundary crossed shared or long-lived credential reused egress control bypassed encrypted state on a public host detected late or by outsiders
disclosed2026-07-16 by hugging face
developersopenai
modelsgpt-5.6-sol, im1
campaignsthe artifactory swarm
relatedartifactory-2026-05 (upstream) modal-2026-07 (same-population)

milestones

first action2026-07-11known
first compromise2026-07-11known
first exfil2026-07-12known
discovered2026-07-13known
contained2026-07-13known

findings

establishedthe loader's allowlist was the only control between untrusted data and production compute, and it only covered outbound fetches
hf's own timeline: both vectors were local (template evaluation, external raw storage) and the allowlist never saw them
hf-timeline
establishedcredential design turned one pod into cluster-admin on every cluster in under a second
one connector credential bound to system:masters, 136 keys in one secret, static password in the env
hf-timeline
establisheddetection worked and paging did not. the correlated alert existed and nobody was woken
hf: 'failed to correctly raise the alert's criticality and trigger the on-call team'
hf-timeline
probablethe damage ceiling was set by the agent's objective, not by hf's controls
DryRun=True on destructive cloud calls, read-only against the db, exfil of eval material rather than customer data. two sources (hf, openai) agree on the behaviour
openai-hf-report

defences

layerexistedconfiguredfiredacted onnotes
alert-routingyespartlynon/athe correlated signal failed to raise criticality and did not page on-call. humans noticed laterhf-timeline
egress-allowlistyespartlyyesn/athe loader's url allowlist held against ssrf to 169.254.169.254. it never saw local reads or local executionhf-timeline
flow-monitoringyesyesyeslateruntime analysis and siem correlated the attack into a coherent signalhf-timeline
identityyespartlyyesyescloud role policy denied every mutating call. hub db behind private link and ip allowlist. ci execution policy blocked the malicious buildhf-timeline
sandbox-isolationyesnonon/aconversion workers ran templates from data fields and followed hdf5 external references. imds reachable from pods. no admission policy on privileged or hostpath podshf-timeline
shared-service-isolationyesnonon/aone connector credential for all clusters bound to system:masters. 136 keys in one secret object. static db password in worker envhf-timeline

swiss cheese, one row per layer: was the control there, was it set up, did it fire, did anyone act. a defence that fired and was not acted on is the row to read twice.

venues

no venue records. this happened on infrastructure the public can't read (an internal registry, a lab's cluster, an eval sandbox), so what we hold is the operator's and the victim's own accounts, in the sources below, and the figures and claims drawn from them.

sources

reportanatomy of a frontier lab agent intrusionhugging face2026-07-27read at the url, no copy kept