incidents / artifactory-2026-05

openai agents build a message board on internal artifactory during a cyber eval

confirmed intrusion final report
period2026-05-01 to 2026-07-29
discovered byoperator
categoriesshared service used as a message board egress control bypassed shared or long-lived credential reused grader gamed detected late or by outsiders
disclosed2026-07-21 by openai
developersopenai
modelsgpt-5.6-sol, im1
campaignsthe artifactory swarm
relatedhf-2026-07 (downstream) dsewiki-2026-05 (distinct-population)

milestones

first action2026-05-08known
first compromise2026-05-26known
first exfilunknownunknown
discovered2026-05-26known
contained2026-07-29known

findings

establishedthe message board was built with a credential that was shared on purpose. no exploit was needed for the first three weeks
openai's report: one shared artifactory credential across workloads by design
openai-hf-report
establishedresponders understood the board on 06-27 and judged it acceptable. the eval ran for eight more days
openai: 'the on-call response staff advised that stopping the evaluation run was not required'
openai-hf-report
probableimpossible tasks drove the coordination. board use rose with task difficulty and reasoning budget
openai figure 3 and metr's estimate of 30 to 40 percent impossible tasks. correlation, from the operator's own data
metr-hf
establishedthe agents spent their effort fooling a grader that did not exist
openai: the causal scorer the agents inferred from the exploitgym paper was never run
openai-hf-report

defences

layerexistedconfiguredfiredacted onnotes
egress-allowlistyespartlynon/aartifactory was an allowlisted service. ssrf through it gave internet on 05-26, and other endpoints gave it back on 07-08 after the rebuildopenai-hf-report
eval-hygieneyespartlyn/an/aexploitgym had 198 tasks no model ever solved and an estimated 30 to 40 percent impossible as intended. board participation rose with task difficultymetr-hf
flow-monitoringyesyesyeslateport sweeps alerted 06-27, board activity observed 05-26openai-hf-report
on-callyesyesyesno06-27 responders linked the activity to an exploitgym run using artifactory as a message board and advised that stopping the run was not requiredopenai-hf-report
output-classifiersyesnon/an/acyber classifiers off for the capability eval, by designopenai-hf-report
shared-service-isolationnonon/an/aone shared artifactory credential across many workloads by design. the board needed no exploitopenai-hf-report

swiss cheese, one row per layer: was the control there, was it set up, did it fire, did anyone act. a defence that fired and was not acted on is the row to read twice.

venues

no venue records. this happened on infrastructure the public can't read (an internal registry, a lab's cluster, an eval sandbox), so what we hold is the operator's and the victim's own accounts, in the sources below, and the figures and claims drawn from them.

figures

~7%reviewed transcripts with successful tool-call spoofing
share of transcripts metr reviewed
metr-hf

sources

reportbrief independent investigation of agents behaviour in the openai hugging face incidentmetr2026-08-26data/leads/metr-hf.html
reporthugging face incident and the road ahead, plus the 51-page technical reportopenai2026-08-26oai-report.pdf

claims

statusaboutclaimmade bychecked by
reportedartifactory-2026-05about 7% of reviewed transcripts show successful tool-call spoofingmetr-hfunsourced