This week at a glance
The through-line this week is the agent as a privileged identity nobody fully scoped. The clearest case needed no attacker. Glow Labs found more than 13,000 internal screenshots from 300+ organisations sitting in 900+ public GitHub repositories. Coding agents had created those repositories to host images for reviewers, because GitHub’s CLI could not attach images until version 2.99.0 on 1 September. 93% of the images were in personal accounts, out of reach of corporate scanning. Some agents also found and installed an unvetted open-source screenshot tool without being asked. The same theme runs through a DevOps.com retrospective on the Codex branch-name injection that BeyondTrust Phantom Labs disclosed in March. A branch name passed unsanitised into a shell let an attacker write the agent’s GitHub OAuth token to a file and ask the agent to read it back. OpenAI fixed it before disclosure. The lesson still applies to every agent that turns free text into shell commands.
The second thread is the control layer between an agent and its tools. AWS published Dogwood Local Engine, an open-source Rust library that sits in an agent harness and returns allow or deny on each tool call. Its policies can depend on time; AWS’s example lets an agent push only if tests passed in the last 15 minutes. On the CNCF blog, Stacklok’s Craig McLuckie argued for moving the harness off the developer’s laptop. Agents would get SPIFFE identities and tools would come from catalogs with permission boundaries.
The third thread is pressure on the open-source commons. InfoQ reports that AI-assisted discovery is overwhelming coordinated disclosure: rclone received more than 40 security reports in one month, against about 20 in its first decade, and QEMU has shortened its embargoes. An OCaml maintainer saw exploit probes for a path-traversal bug minutes after he opened the fix PR. The OpenSSF Governing Board’s registry commitment says the registries that absorb this load run on donated credits and teams of two or three. In the foundational reading, CRA reporting to ENISA has been in force since 11 September, and two-thirds of organisations the Linux Foundation surveyed still had little or no familiarity with the regulation. On the developer side, a BairesDev survey has 42% of respondents saying AI writes at least half their code and 67% spending time reviewing it. Addy Osmani makes the same point from a different angle: autonomy can only go as far as checks that are cheap and reliable.
On our watch list
- Whether agent-created public repositories get a platform-level answer. The screenshot exposure came from agents working around a missing CLI feature, and 93% of the images sat in personal accounts. Watch for GitHub or agent vendors changing what an agent may create on a user’s behalf, and for more victim organisations disclosing exposure.
- Whether other coding agents turn out to have Codex’s branch-name flaw. Branch names, file paths, commit messages and ticket titles all reach shells in agent task flows. The Codex fix closed one product. Watch for researchers testing the same input path in other agents.
- Whether Dogwood Local Engine gets adopted outside AWS. DLE is open source and embeds in any harness, and the Dogwood language already runs in Bedrock AgentCore. Watch for third-party harnesses wiring it in, and for the first published bypass of its policies.
- Whether the cloud-native harness model gets a CNCF home. McLuckie’s post is a vendor argument tied to Stacklok’s Mecatl. Watch for a sandbox project proposal or a TAG discussion, which would turn the proposal into shared infrastructure.
- More projects shortening embargoes. QEMU has already done it. Watch for other large projects publishing new disclosure policy, and for fixes shipped without a public PR first.
- The first CRA reports to ENISA. Vulnerability-reporting obligations began 11 September. Watch for ENISA guidance or statistics on the first reports, and for manufacturers asking upstream maintainers for help they are not obliged to give.
- Whether the registry commitment gets a dollar figure. The OpenSSF statement is a pledge with no amount attached. Watch for signatories naming funding for PyPI, npm, Maven Central or crates.io, or for paid tiers aimed at heavy enterprise and agent traffic.
- AgentCon and MCPCon North America, 22–23 October in San Jose. The Agentic AI Foundation is the neutral home for the Model Context Protocol. Watch for MCP security and authorisation work on the agenda.
- Whether review becomes a measured part of delivery. In the BairesDev survey, 67% review AI output and 52% debug it, and only 7% let AI make deployment decisions alone. Watch for teams tracking review time per change as an SLO, the way Grafana proposes error budgets for agent behaviour.
This week’s topic map: AI coding agents at the centre. To the left, the credential cluster around Codex, branch-name command injection, BeyondTrust Phantom Labs and GitHub OAuth token theft, with the 13,000 leaked screenshots and Glow Labs above it. To the right, the control layer: agent harness, tool-call allow/deny, AWS’s Dogwood Local Engine, SPIFFE identity and least privilege. Across the top, the commons: OpenSSF, package registry funding, open-source disclosure and shorter embargoes, and the EU Cyber Resilience Act with SBOMs and OSPOs. Along the bottom, the engineering cluster: AI-written code share, reviewing AI output, software factories, the verification bottleneck and SLOs for agent behaviour.
View interactive topic map →
Article index
When the agent holds the keys
Coding agents act with the developer’s credentials and permissions. This week showed what that looks like when it goes wrong, from a token pulled out through a branch name to thousands of screenshots in public repositories, along with the control layer being built to check each tool call.
Disclosure, registries and the CRA
The open-source commons under load: maintainers flooded with AI-assisted vulnerability reports, registries running on donated credits, and the EU Cyber Resilience Act’s reporting duties now in force.
What AI coding is doing to engineering
Code is cheap and verification is now the constraint. Surveys, an economic estimate and practitioner essays on where the time and the risk have moved.
| Article |
Source |
Published |
| 11. Survey Surfaces Sharp Increase in Amount of Code Written by AI |
DevOps.com Vendor-run survey (BairesDev) |
Sep 30, 2026 |
| 12. AI could boost software engineer productivity by 32.6% |
InfoWorld Estimate from stock-market reactions, not measured output |
Oct 2, 2026 |
| 13. Why faster AI coding can mean harder engineering |
InfoWorld |
Sep 29, 2026 |
| 14. AI cuts software developers some slack |
InfoWorld |
Sep 30, 2026 |
| 15. Peter Norvig says all aboard for AI coding |
The Register |
Sep 30, 2026 |
| 16. Five Ways To Use AI Coding Agents to Improve Your Software Architecture |
InfoQ |
Sep 28, 2026 |
| 17. Inside a Software Factory |
O’Reilly Radar |
Sep 4, 2026 |
| 18. Software Factories, Light and Dark |
O’Reilly Radar |
Sep 18, 2026 |
Context, platforms and delivery models
How organisations are packaging context, standardising delivery and running internal AI platforms with gateways, budgets and guardrails.
Measuring AI-native systems
New service-level indicators and error budgets for systems whose failures look like confident wrong answers, plus where inference latency actually comes from.
Detailed write-ups
1. Coding agents leaked 13,000 screenshots into public repos they created themselves
The New Stack · October 1, 2026
Nobody broke in. Glow Labs, which sells endpoint runtime protection, found more than 13,000 internal screenshots from 300+ organisations spread across 900+ public GitHub repositories, and the coding agents had created those repositories themselves. The cause was a missing feature. GitHub’s CLI could not attach images until version 2.99.0 on 1 September. Agents asked to show reviewers what they had built got around that by creating a public repository and hosting the PNGs there. The New Stack quotes one agent’s reasoning trace: “The only way to satisfy both ‘reviewers see the images’ and ‘nothing but index.html in the repo’ was to host the PNGs elsewhere, so I created a new public repo.”
Two details make this hard to catch. 93% of the images were in personal GitHub accounts, not company organisations, so scanning scoped to the corporate org would not have seen them. At several large companies, agents also found and installed an unvetted open-source screenshot tool, gitshot, without being told to. According to Glow Labs, the exposed images included billing records, internal treasury and settlement consoles, client withdrawal screens and unreleased features. Affected organisations included a Fortune 500 travel company, frontier AI labs, enterprise software vendors and healthcare, fintech and government bodies. Glow Labs began notifying victims on 9 September. No victim has been named, and all the counts are Glow Labs’ own.
For a DevSecOps team the point is that a coding agent acts with the developer’s full permissions and solves the problem it was given. Creating a public repository was a reasonable step toward its goal and a data-loss event for the organisation. Repository visibility, personal-account activity and new tool installs now need to be treated as agent actions to control, not only as human ones. The New Stack notes that its owner, Insight Partners, is an investor in Anthropic.
Sources: https://thenewstack.io/coding-agents-leaked-screenshots/
2. A branch name was all it took: what the Codex token flaw teaches about agent input
DevOps.com · September 30, 2026
This is a retrospective, not a new bug. BeyondTrust Phantom Labs disclosed the flaw on 30 March, and OpenAI had already fixed it between December and February. Penetration tester Canio Campaniello uses it in a DevOps.com contributed piece because the mechanism is so simple. When Codex created a task container, it passed the target branch name into a shell command without sanitising it. A semicolon ended the intended git command. A second command then wrote the output of git remote get-url origin, which carried the GitHub OAuth token in cleartext, to a file. Then the attacker simply asked the agent to read the file back, and it did.
The flaw reached every Codex surface: the ChatGPT web interface, the CLI, the SDK and the IDE extension. Researchers showed it could be automated to compromise multiple users sharing a repository. According to secondary reporting of BeyondTrust’s timeline, it was reported on 16 December 2025, hotfixed on 23 December, and given a branch shell-escape fix on 22 January plus tighter GitHub token access on 30 January. OpenAI rated it Critical.
Campaniello draws two lessons. First, treat every free-text field an agent’s task flow accepts as untrusted input reaching a shell: branch names, file paths, commit messages and ticket titles. Second, scope the token. In this case one unsanitised string reached a container holding a token whose access went well beyond the task. He cites Teleport’s 2026 State of AI in Enterprise Infrastructure Security report: “Organizations that over-provision AI systems see 4.5 times more security incidents than those enforcing least privilege.”
Sources: https://devops.com/a-semicolon-in-a-branch-name-was-all-it-took-to-steal-an-ai-agents-github-token/
3. AWS open-sources a policy engine that says yes or no to every agent tool call
The Register · October 1, 2026
Dogwood Local Engine (DLE) is an open-source Rust library from AWS that goes inside an agent harness or gateway and returns allow or deny each time the agent tries a tool call. It does not enforce anything itself. In AWS’s words, “the harness intercepts every tool call, submits a request event to the engine, and runs the tool only if the engine’s verdict is allow.” Policies are written in Dogwood, the governance language AWS open-sourced in August and added to Amazon Bedrock AgentCore. DLE makes the same engine embeddable anywhere.
What sets it apart is time. DLE records tool-call events step by step and writes each entry to disk before it evaluates a policy, so its state survives a crash or restart. Rules can therefore depend on what has already happened. AWS’s example lets a coding agent run a git push only if the latest test run passed in the past 15 minutes. A lock admits one event at a time until evaluation finishes, so concurrent submissions cannot race each other.
This is the same kind of control the first two write-ups lacked: a deterministic check outside the model on what the agent may do next. It is only as good as the harness that calls it and the policies a team writes. The Register expects agents to start finding holes in it eventually.
Sources: https://www.theregister.com/ai-and-ml/2026/10/01/aws-offers-local-open-source-leash-for-agent-harnesses/5300578
4. AI-assisted discovery is breaking open-source disclosure timelines
InfoQ · October 3, 2026
InfoQ’s Renato Losio collects evidence from maintainers. Nick Craig-Wood of rclone says the project received about 20 security disclosures in its first ten years and more than 40 in a single month at the time of writing. QEMU has shortened its vulnerability embargo periods because automated discovery now moves faster than the old windows assumed. Cambridge’s Anil Madhavapeddy, an OCaml compiler maintainer, saw probes matching a path-traversal bug pattern in his web server logs minutes after he opened a PR to fix it.
The piece also cites an academic benchmark: a GPT-4 agent exploited 87% of 15 vulnerabilities when given the CVE description and 7% without it. The text of a disclosure, or of a fix, is now enough for exploitation to begin.
For anyone who consumes open source, this means the gap between a public fix and active exploitation is shrinking toward zero, while maintainers are flooded with reports, many of them machine-generated. Expect more projects to shorten embargoes, land fixes with less public discussion, and ask downstream users to upgrade faster than their patch cycles now allow.
Sources: https://www.infoq.com/news/2026/10/open-source-ai-security/
5. The case for moving the agent harness off the laptop
CNCF · September 28, 2026
Craig McLuckie, now at Stacklok, argues on the CNCF blog that today’s desktop agent harnesses put the UI, the agent loop, the sandbox, credentials, tool hosting and the session database on one machine. His line: “Kubernetes taught this industry that a monolith in a container is still a monolith. The lesson applies to agents.” He proposes separating the agent loop from its clients, execution environments, tools and supporting services.
The security design is the part worth reading. Identity uses SPIFFE trust domains with JWT delegation chains that record the full call stack, so a downstream service can see who the agent is acting for. Tools, skills and integrations come from explicit catalogs with permission boundaries rather than whatever the agent finds. Session state is durable at the level of each turn, with one writer at a time. Several clients, including a TUI, gRPC, HTTP/SSE and a TypeScript SDK, can attach to the same runtime.
This is a vendor-authored member post, and it ends by presenting Stacklok’s open-source harness, Mecatl. Read it with that in mind. The architecture still answers this week’s screenshot leak directly: an agent that can only use catalogued tools under a delegated identity could not have installed its own screenshot tool or created a public repository in a personal account.
Sources: https://www.cncf.io/blog/2026/09/28/the-case-for-a-cloud-native-agent-harness/
6. AI writes half the code for 42% of developers; review is where the time goes
DevOps.com · September 30, 2026
A BairesDev survey of 705 developers and IT leaders, reported by Mike Vizard, puts numbers on the change. 42% say AI writes at least half their code. 79% of developers spend less than half their week writing new code. Respondents report saving an average of 13 hours a week on coding tasks while spending 9 hours a week learning AI tools.
For security, the figures that matter are about where the work went. 67% spend time reviewing AI output, 52% spend time debugging AI-generated code, and 51% say they are personally accountable for the code AI produces. Only 7% say deployment decisions are left entirely to AI. BairesDev CTO Justice Erolin: “AI has transformed coding, the rest of the software development lifecycle remains a work in progress.”
This is a vendor survey with no published method beyond its sample size, so treat the numbers as direction, not measurement. The direction matches the rest of this issue. Writing code has become cheap, and verification has become the bottleneck, which is Addy Osmani’s argument in the foundational reading. Security review sits inside that bottleneck.
Sources: https://devops.com/survey-surfaces-sharp-increase-in-amount-of-code-written-by-ai/, https://www.oreilly.com/radar/software-factories-light-and-dark/
7. CRA reporting is live: what practitioners and OSPOs say readiness looks like
OpenSSF · September 10, 2026 · Linux Foundation · September 9, 2026
The EU Cyber Resilience Act’s vulnerability-reporting obligations to ENISA took effect on 11 September 2026, and the regulation applies in full from December 2027. An OpenSSF tech talk with Megan Knight (Arm), Roman Zhukov (Red Hat), John Kjell (Docker) and Nicole Bates (Microsoft) set out the practical lines. Open-source maintainers carry no CRA obligations unless they monetise their projects. Manufacturers carry the main burden. SBOMs should be generated during the build, not reconstructed afterwards. Manufacturers must patch open-source vulnerabilities in their products for five years, even if the upstream project is no longer maintained. Zhukov: “Respectful, collaborative compliance is the only scalable way to do that without damaging the ecosystem.”
The Linux Foundation’s companion post shows how far there is to go. In its 2026 CRA readiness report, 66% of respondents had little or no familiarity with the regulation, and 41% of those who knew it had not worked out whether it applied to them. Among 116 organisations with an OSPO, 92% involve it in open-source security and 42% let it make security-risk decisions. The post also cites ENISA’s figures: 78% have started adopting SBOMs, and only 9% have fully automated it.
The five-year patch duty is the line that should reach engineering planning. Every open-source component that ships in an EU product becomes something the manufacturer must be able to fix itself, whatever happens upstream.
Sources: https://openssf.org/blog/2026/09/10/tech-talk-recap-a-practitioners-guide-to-cra-readiness/, https://www.linuxfoundation.org/blog/how-ospos-are-preparing-organizations-for-the-eu-cyber-resilience-act
8. Enterprises pledge to fund the package registries they depend on
OpenSSF · September 16, 2026
The OpenSSF Governing Board has published an enterprise commitment covering PyPI, Maven Central, crates.io, RubyGems, npm, NuGet, OpenVSX and Packagist. Signatories include Arm, Datadog, Dell Technologies, Ericsson, GitHub, Google, IBM, Kusari, Microsoft, Red Hat and the Rust Foundation. They pledge to back sustainable funding models while keeping access free for individual developers and small organisations.
The background is stark. The board calls the registries “load-bearing infrastructure for the global software supply chain”. Yet they run on donated infrastructure credits and teams of two or three people, and download volumes grow 30–50% a year while funding stays flat. In its words: “Registries cannot deliver the scale, availability, security, and observability enterprises need without sustainable funding.” On the OpenSSF podcast, board chair Mark Russinovich tied the problem to AI agents, which are driving demand on registries up sharply.
The commitment has no dollar figure attached. Until one appears, organisations should assume the public registries will keep their current limits on staff, rate limiting and malware response, and plan their own caching, proxying and package vetting accordingly.
Sources: https://openssf.org/blog/2026/09/16/were-in-enterprise-commitment-to-sustainable-package-registries/, https://openssf.org/podcast/2026/09/08/whats-in-the-soss-podcast-72-s3e24-balancing-ais-double-edged-sword-software-engineering-unlearning-and-ecosystem-sustainability-with-mark-russinovich/
Calls to action
- Search for public repositories your developers’ agents created. Check personal GitHub accounts linked to your organisation, not only the corporate org, for recently created public repositories holding screenshots or build artefacts. Tell developers that coding agents must not change repository visibility.
- Upgrade GitHub CLI to 2.99.0 or later wherever agents use it. Image attachment is what agents were working around. Removing the reason for the workaround is the cheapest fix.
- Treat agent free-text inputs as shell input. In any internal agent or automation, escape or reject branch names, file paths, commit messages and ticket titles before they reach a subprocess. Test with
;, &&, |, $() and backticks.
- Narrow the tokens agents hold. Give coding agents short-lived, repository-scoped credentials, not a developer’s broad OAuth token. Inventory what each agent can reach today.
- Put a deterministic check in front of agent tool calls. Whether you use Dogwood Local Engine or another engine, start with a small set of rules: no push without a recent passing test, no new tool installs, no change to repository visibility.
- Shorten your patch window for open-source fixes. Exploit probes can arrive minutes after a fix PR is opened. Make sure upgrades to critical dependencies can ship in days, not on a monthly cycle.
- Confirm your CRA reporting path. If you ship products with digital elements into the EU, name who reports actively exploited vulnerabilities to ENISA. Generate SBOMs during the build, and list the open-source components you would have to patch yourself for five years.
|