{"id":5781,"date":"2026-08-30T15:14:58","date_gmt":"2026-08-30T20:14:58","guid":{"rendered":"https:\/\/www.cybersecurityinstitute.com\/blog\/?p=5781"},"modified":"2026-08-30T15:14:58","modified_gmt":"2026-08-30T20:14:58","slug":"ai-ops-weekly-august-30-2026","status":"publish","type":"post","link":"https:\/\/www.cybersecurityinstitute.com\/blog\/?p=5781","title":{"rendered":"AI Ops Weekly &mdash; August 30, 2026"},"content":{"rendered":"<style>\n.single .entry-title,\n.single .entry-header .entry-title,\n.single .post-title,\n.single header.entry-header h1,\n.single h1.entry-title,\n.single .page-title,\n.post-template-default h1.entry-title,\n.post-template-default .entry-header,\narticle .entry-header,\narticle .entry-title { display: none !important; }\n.single .entry-header { margin: 0 !important; padding: 0 !important; }\n.single .entry-content { margin-top: 0 !important; padding-top: 0 !important; }\n<\/style>\n<table role=\"presentation\" class=\"wrapper\" cellpadding=\"0\" cellspacing=\"0\" border=\"0\" width=\"100%\">\n<tr>\n<td align=\"center\">\n<table role=\"presentation\" class=\"container\" cellpadding=\"0\" cellspacing=\"0\" border=\"0\" width=\"680\">\n<p>        <!-- Banner --><\/p>\n<tr>\n<td class=\"banner\" style=\"background-color:#0e7490;background:linear-gradient(135deg,#0e7490 0%,#0891b2 100%);padding:36px 32px;color:#ffffff;\">\n<p class=\"date\" style=\"color:#ffffff !important;\">August 30, 2026 &middot; Weekly Edition<\/p>\n<h1 style=\"color:#ffffff !important;\">AI Ops<\/h1>\n<p class=\"tagline\" style=\"color:#ffffff !important;\">This week the beat narrowed to a single question: what does it actually cost to run AI in production, and can you see it? Elastic closed its acquisition of Deductive AI and bought itself an incident-reasoning layer; Grafana said AI demand pushed it past $600M ARR; Dynatrace published a state-of-SRE report on where enterprise AI is breaking. Underneath the vendor news sits the engineering: telemetry pipelines as a cost-control tier, agent traces arriving as application data, NVIDIA shipping shadow-engine recovery for lost inference capacity, Google Cloud writing down dynamic capacity management and gVisor-sandboxed Ray, and Kubernetes 1.37 clearing out kube-dns, IPVS and cgroup v1. Twenty-four stories.<\/p>\n<\/td>\n<\/tr>\n<p>        <!-- At a glance --><\/p>\n<tr>\n<td class=\"content\">\n<h2>This week at a glance<\/h2>\n<p>The spine of the week is <strong>the cost and observability of running AI in production<\/strong> &mdash; and the two are the same problem viewed from different ends. The New Stack put it most directly: observability already has a data problem, and AI is about to make it worse. Observability already runs at <strong>10&ndash;30% of infrastructure spend<\/strong> while returning only partial access to the data, which is why every standard coping mechanism is subtractive &mdash; sample the traces, index less of the logs, ratchet retention from 30 days to seven to three. Agent workloads make that arithmetic worse, and the survey numbers are stark: <strong>54%<\/strong> of enterprises saw telemetry volume triple in a year, with <strong>43%<\/strong> of the growth coming from AI\/ML workloads; average observability spend now sits at <strong>$3.17 million<\/strong> and is rising <strong>28%<\/strong> year over year; and <strong>59%<\/strong> of organisations have already terminated or delayed an agentic AI deployment because of monitoring cost. The practical consequence is that <strong>telemetry pipelines<\/strong> stop being plumbing and become a budget control: the tier where you decide what is sampled, what is aggregated at the edge, what is routed to cheap object storage and what earns a place in the expensive index. The companion piece on <strong>agent traces becoming application data<\/strong> is the other half of the argument &mdash; at the 500,000 browser events a day one agent platform reports, a trace of an agent run is no longer a debugging artifact you discard after a week, it is the record of what the system decided and why, which means retention, schema and access control now have product requirements attached rather than just ops ones.<\/p>\n<p>The money story ran in parallel and reached the same place. DevOps.com made the sharpest framing of the week with the observation that <strong>an AI coding budget is becoming a variable cloud bill<\/strong> &mdash; <strong>GitHub Copilot<\/strong>&rsquo;s mid-2026 move to usage-based billing turned a predictable per-seat licence into a metered credit pool that premium models and agentic features draw down, so two engineers on the same plan now generate very different costs. InfoWorld went after the same total from the other direction, arguing that the biggest hidden AI cost is <em>not<\/em> GPUs but context bloat &mdash; paying to process noise you never filtered out &mdash; at a moment when <strong>73%<\/strong> of enterprises say their AI costs have already outpaced what they budgeted. Then the harder question: usage is measurable, value is not. IBM finds only <strong>29%<\/strong> of executives can confidently measure AI ROI and just <strong>25%<\/strong> of AI initiatives deliver the return expected of them. That gap is where <strong>Tempo<\/strong> planted its flag this week, launching Workforce Intelligence to tie AI spend to Jira work items &mdash; an attempt to make the denominator of the ROI fraction something a delivery organisation already tracks. Whether or not Tempo&rsquo;s particular join is the right one, the pattern is worth noting: the market is moving from &ldquo;how many tokens did we burn&rdquo; to &ldquo;which unit of delivered work did we burn them on.&rdquo;<\/p>\n<p>The consolidation story is the third strand, and it is the one with the clearest strategic read. <strong>Elastic completed its acquisition of Deductive AI<\/strong> on August 24, on undisclosed terms, folding in an AI SRE agent that forms and tests hypotheses across code, telemetry and organisational knowledge to reach a root cause &mdash; CEO Ash Kulkarni&rsquo;s framing is that engineering teams today are &ldquo;drowning in telemetry but starved for answers.&rdquo; The logic is the same logic every observability vendor is now executing against: the differentiator is no longer ingest or storage, it is whether the platform can reason over what it has stored and shorten the distance between a page and an explanation. <strong>Grafana Labs<\/strong> supplied the demand-side evidence, telling Business Insider that ARR reached <strong>$600M<\/strong> with AI a material driver &mdash; more AI systems in production means more signals, more dashboards and more people who need to understand a stack they did not write. And <strong>Dynatrace<\/strong>&rsquo;s State of SRE and Platform Engineering 2026, drawn from 919 senior IT leaders at enterprises above $500M in revenue, found <strong>67%<\/strong> of SREs now name AI model monitoring their top use case. Read the three together and the market thesis is unambiguous &mdash; observability is being repriced as the control plane for AI operations, and the vendors are buying, hiring and reporting accordingly.<\/p>\n<p>Beneath all of that, the infrastructure engineering was unusually concrete. <strong>NVIDIA<\/strong> documented shadow engine recovery in <strong>Dynamo<\/strong>, cutting recovery of a failed inference worker from <strong>283 seconds<\/strong> to <strong>7.3<\/strong> in its own benchmark &mdash; a direct answer to the fact that an inference fleet degrades differently from a stateless web tier, because the expensive state is resident in the accelerator. <strong>Google Cloud<\/strong> published two pieces that belong on a platform team&rsquo;s reading list: best practices for dynamic capacity management, which is what you need when only <strong>17%<\/strong> of IT leaders believe their infrastructure can carry the agent workloads <strong>90%<\/strong> of enterprises intend to deploy within three years, and gVisor sandboxes for distributed <strong>Ray<\/strong> clusters on <strong>GKE<\/strong>, scaled with Anyscale to <strong>100,000 sandboxes in 17.3 seconds<\/strong>. The New Stack asked why real-time AI at scale is so hard and answered in tail latency rather than throughput &mdash; p99 under 10 ms at a concurrency of 300, three seconds at around 740,000 operations per second. Around the edges: <strong>DevOps.com<\/strong> on durable execution as the missing runtime for long-running agents, The New Stack on the three roles agents play in a developer platform, <strong>Nutanix<\/strong> shipping Enterprise AI 2.8 and an Agent Gateway, <strong>neoclouds<\/strong> chasing 20% of a $267 billion AI cloud market by 2030, and <strong>Kubernetes 1.37<\/strong> starting the clock on kube-dns, IPVS and cgroup v1 &mdash; the unglamorous maintenance work that decides whether next year&rsquo;s upgrade is routine or an outage.<\/p>\n<p>            <!-- Topic map --><\/p>\n<div class=\"topic-map\">\n              <img decoding=\"async\" src=\"https:\/\/www.cybersecurityinstitute.com\/blog\/wp-content\/uploads\/2026\/08\/topic-map-aiops-2026-08-30.png\" alt=\"Topic map of this week's AI Ops themes: AI cost management at the centre, linked to token economics, hidden AI costs, proving AI value and Tempo; a telemetry hub linking OpenTelemetry, observability data volume and agent traces; an observability consolidation hub linking Elastic, Deductive AI, Grafana Labs and Dynatrace to incident AI; an inference and capacity hub linking NVIDIA Dynamo, Google Cloud, neoclouds and real-time AI; and a platform hub linking agent runtimes, workload sandboxing with gVisor and Ray on GKE, Nutanix and Kubernetes 1.37 deprecations\" loading=\"eager\"><\/p>\n<p class=\"caption\">This week&rsquo;s topic map &mdash; AI cost management sits at the centre, wired to token economics, hidden costs beyond GPUs, the problem of proving AI value, and Tempo&rsquo;s attempt to tie spend to delivered work. A telemetry hub connects OpenTelemetry, observability data volume and agent traces as application data; an observability consolidation hub runs from Elastic and Deductive AI through Grafana and Dynatrace into incident AI; an inference and capacity hub ties NVIDIA Dynamo, Google Cloud, neoclouds and real-time serving together; and a platform hub links agent runtimes, gVisor-sandboxed Ray on GKE, Nutanix and the Kubernetes 1.37 deprecations.<\/p>\n<p>              <!-- INTERACTIVE_MAP_LINK_START --><\/p>\n<p style=\"margin:10px 0 0;text-align:center;\"><a href=\"https:\/\/www.cybersecurityinstitute.com\/blog\/?p=5780\" target=\"_blank\" rel=\"noopener\" style=\"display:inline-block;padding:8px 18px;background-color:#0f172a;color:#ffffff !important;text-decoration:none;border-radius:6px;font-size:13px;font-weight:600;\">View interactive topic map &rarr;<\/a><\/p>\n<p><!-- INTERACTIVE_MAP_LINK_END -->\n            <\/div>\n<p>            <!-- Article index --><\/p>\n<h2>Article index<\/h2>\n<p style=\"font-size:13px;color:#6b7280;font-style:italic;margin:0 0 6px 0;\">24 articles, grouped by sub-theme. &ldquo;Weekly News&rdquo; = this week&rsquo;s coverage window (August 23&ndash;30); &ldquo;Foundational Reading&rdquo; = longer-form reference reading on the beat.<\/p>\n<h3>Weekly News<\/h3>\n<h4>Consolidation and the observability data problem<\/h4>\n<div class=\"cluster-intro\">Elastic buys a reasoning layer, Grafana reports what AI demand is doing to its top line, and two pieces explain why the data underneath all of it is getting harder to afford.<\/div>\n<table class=\"index-table\">\n<tr>\n<th>Article<\/th>\n<th>Source<\/th>\n<th>Published<\/th>\n<\/tr>\n<tr>\n<td>1. <a href=\"https:\/\/www.businesswire.com\/news\/home\/20260824863894\/en\/Elastic-Completes-Acquisition-of-Deductive-AI\">Elastic Completes Acquisition of Deductive AI<\/a><\/td>\n<td class=\"src\">Business Wire<\/td>\n<td class=\"dt\">Aug 24, 2026<\/td>\n<\/tr>\n<tr>\n<td>2. <a href=\"https:\/\/thenewstack.io\/opentelemetry-observability-telemetry-storage\/\">Observability has a data problem. AI is about to make it worse.<\/a><\/td>\n<td class=\"src\">The New Stack<\/td>\n<td class=\"dt\">Aug 26, 2026<\/td>\n<\/tr>\n<tr>\n<td>3. <a href=\"https:\/\/www.businessinsider.com\/grafana-labs-arr-600-million-ai-demand-2026-8\">Grafana&rsquo;s ARR Hits $600M As AI Boosts Demand, CEO Says<\/a><\/td>\n<td class=\"src\">Business Insider<\/td>\n<td class=\"dt\">Aug 26, 2026<\/td>\n<\/tr>\n<tr>\n<td>4. <a href=\"https:\/\/www.dynatrace.com\/news\/press-release\/state-of-sre-platform-engineering-2026\/\">As AI Scales Across Enterprises, Breaking Points Emerge (State of SRE and Platform Engineering 2026)<\/a><\/td>\n<td class=\"src\">Dynatrace<\/td>\n<td class=\"dt\">Aug 25, 2026<\/td>\n<\/tr>\n<\/table>\n<h4>Telemetry pipelines and the cost of agent data<\/h4>\n<div class=\"cluster-intro\">Where the spend actually accrues: the pipeline tier that decides what gets kept, the traces agents emit, and two hard looks at what an AI line item does to a budget.<\/div>\n<table class=\"index-table\">\n<tr>\n<th>Article<\/th>\n<th>Source<\/th>\n<th>Published<\/th>\n<\/tr>\n<tr>\n<td>5. <a href=\"https:\/\/thenewstack.io\/agentic-ai-telemetry-costs\/\">How telemetry pipelines keep AI agent costs under control<\/a><\/td>\n<td class=\"src\">The New Stack<\/td>\n<td class=\"dt\">Aug 25, 2026<\/td>\n<\/tr>\n<tr>\n<td>6. <a href=\"https:\/\/thenewstack.io\/agent-traces-application-data\/\">When AI agent traces become application data<\/a><\/td>\n<td class=\"src\">The New Stack<\/td>\n<td class=\"dt\">Aug 26, 2026<\/td>\n<\/tr>\n<tr>\n<td>7. <a href=\"https:\/\/devops.com\/your-ai-coding-budget-is-becoming-a-variable-cloud-bill\/\">Your AI Coding Budget Is Becoming a Variable Cloud Bill<\/a><\/td>\n<td class=\"src\">DevOps.com<\/td>\n<td class=\"dt\">Aug 26, 2026<\/td>\n<\/tr>\n<tr>\n<td>8. <a href=\"https:\/\/www.infoworld.com\/article\/4210670\/why-your-biggest-hidden-ai-cost-isnt-gpus.html\">Why your biggest hidden AI cost isn&rsquo;t GPUs<\/a><\/td>\n<td class=\"src\">InfoWorld<\/td>\n<td class=\"dt\">Aug 24, 2026<\/td>\n<\/tr>\n<\/table>\n<h4>Cost accountability and the compute market<\/h4>\n<div class=\"cluster-intro\">Measuring usage turns out to be the easy half. Proving value, attributing spend to delivered work, and who ends up supplying the compute.<\/div>\n<table class=\"index-table\">\n<tr>\n<th>Article<\/th>\n<th>Source<\/th>\n<th>Published<\/th>\n<\/tr>\n<tr>\n<td>9. <a href=\"https:\/\/www.infoworld.com\/article\/4213146\/enterprises-can-measure-ai-usage-but-the-hard-part-is-proving-that-it-actually-delivered-value.html\">Enterprises can measure AI usage, but the hard part is proving that it actually delivered value<\/a><\/td>\n<td class=\"src\">InfoWorld<\/td>\n<td class=\"dt\">Aug 25, 2026<\/td>\n<\/tr>\n<tr>\n<td>10. <a href=\"https:\/\/siliconangle.com\/2026\/08\/25\/tempo-launches-workforce-intelligence-to-tie-ai-spend-to-jira-work-items\/\">Tempo launches Workforce Intelligence to tie AI spend to Jira work items<\/a><\/td>\n<td class=\"src\">SiliconANGLE<\/td>\n<td class=\"dt\">Aug 25, 2026<\/td>\n<\/tr>\n<tr>\n<td>11. <a href=\"https:\/\/www.infoworld.com\/article\/4211105\/can-neoclouds-corner-ai-compute.html\">Can neoclouds corner AI compute?<\/a><\/td>\n<td class=\"src\">InfoWorld<\/td>\n<td class=\"dt\">Aug 24, 2026<\/td>\n<\/tr>\n<\/table>\n<h4>Inference, capacity and real-time serving<\/h4>\n<div class=\"cluster-intro\">The engineering underneath the invoice &mdash; recovering lost inference capacity, managing accelerator supply dynamically, sandboxing distributed compute, and why real time stays hard.<\/div>\n<table class=\"index-table\">\n<tr>\n<th>Article<\/th>\n<th>Source<\/th>\n<th>Published<\/th>\n<\/tr>\n<tr>\n<td>12. <a href=\"https:\/\/developer.nvidia.com\/blog\/restore-llm-inference-capacity-in-seconds-with-shadow-engine-recovery-in-nvidia-dynamo\">Restore LLM Inference Capacity in Seconds with Shadow Engine Recovery in NVIDIA Dynamo<\/a><\/td>\n<td class=\"src\">NVIDIA Technical Blog<\/td>\n<td class=\"dt\">Aug 25, 2026<\/td>\n<\/tr>\n<tr>\n<td>13. <a href=\"https:\/\/cloud.google.com\/blog\/topics\/ai-infrastructure\/best-practices-for-dynamic-capacity-management\/\">Best practices for dynamic capacity management<\/a><\/td>\n<td class=\"src\">Google Cloud Blog<\/td>\n<td class=\"dt\">Aug 26, 2026<\/td>\n<\/tr>\n<tr>\n<td>14. <a href=\"https:\/\/thenewstack.io\/real-time-ai-scale\/\">Why real-time AI at scale is so hard<\/a><\/td>\n<td class=\"src\">The New Stack<\/td>\n<td class=\"dt\">Aug 23, 2026<\/td>\n<\/tr>\n<tr>\n<td>15. <a href=\"https:\/\/cloud.google.com\/blog\/products\/containers-kubernetes\/gvisor-sandboxes-for-ray-clusters-on-gke\">Bringing gVisor sandboxes to distributed Ray clusters<\/a><\/td>\n<td class=\"src\">Google Cloud Blog<\/td>\n<td class=\"dt\">Aug 25, 2026<\/td>\n<\/tr>\n<\/table>\n<h4>Agent runtimes and platform operations<\/h4>\n<div class=\"cluster-intro\">What agents run on, what the developer platform owes them, and the maintenance work that keeps the substrate upgradeable.<\/div>\n<table class=\"index-table\">\n<tr>\n<th>Article<\/th>\n<th>Source<\/th>\n<th>Published<\/th>\n<\/tr>\n<tr>\n<td>16. <a href=\"https:\/\/devops.com\/the-missing-runtime-for-long-running-ai-agents\/\">The Missing Runtime for Long-Running AI Agents<\/a><\/td>\n<td class=\"src\">DevOps.com<\/td>\n<td class=\"dt\">Aug 26, 2026<\/td>\n<\/tr>\n<tr>\n<td>17. <a href=\"https:\/\/thenewstack.io\/ai-agent-platform-roles\/\">The 3 roles AI agents play in your developer platform<\/a><\/td>\n<td class=\"src\">The New Stack<\/td>\n<td class=\"dt\">Aug 29, 2026<\/td>\n<\/tr>\n<tr>\n<td>18. <a href=\"https:\/\/www.blocksandfiles.com\/hci\/2026\/08\/26\/nutanix-adds-more-rooms-to-its-agentic-ai-building\/5292580\">Nutanix adds more rooms to its agentic AI building<\/a><\/td>\n<td class=\"src\">Blocks &amp; Files<\/td>\n<td class=\"dt\">Aug 26, 2026<\/td>\n<\/tr>\n<tr>\n<td>19. <a href=\"https:\/\/www.theregister.com\/devops\/2026\/08\/26\/kubernetes-cleans-house-bins-legacy-kube-dns-ipvs-and-cgroup-v1\/5292717\">Kubernetes cleans house, bins legacy kube-dns, IPVS, and cgroup v1<\/a><\/td>\n<td class=\"src\">The Register<\/td>\n<td class=\"dt\">Aug 26, 2026<\/td>\n<\/tr>\n<tr>\n<td>20. <a href=\"https:\/\/devops.com\/automated-diagnosis-isnt-automated-understanding-what-postmortems-teach-us-about-building-trustworthy-incident-ai\/\">Automated Diagnosis Isn&rsquo;t Automated Understanding: What Postmortems Teach Us About Building Trustworthy Incident AI<\/a><\/td>\n<td class=\"src\">DevOps.com<\/td>\n<td class=\"dt\">Aug 26, 2026<\/td>\n<\/tr>\n<\/table>\n<h3>Foundational Reading<\/h3>\n<h4>Token economics and model selection<\/h4>\n<div class=\"cluster-intro\">Three reference reads on the arithmetic behind every architecture decision in this issue: what a token actually costs you, when a smaller model is the correct answer, and how to serve inference sensibly.<\/div>\n<table class=\"index-table\">\n<tr>\n<th>Article<\/th>\n<th>Source<\/th>\n<th>Published<\/th>\n<\/tr>\n<tr>\n<td>21. <a href=\"https:\/\/www.oreilly.com\/radar\/tokens-arent-dollars\/\">Tokens Aren&rsquo;t Dollars<\/a><\/td>\n<td class=\"src\">O&rsquo;Reilly Radar<\/td>\n<td class=\"dt\">Aug 28, 2026<\/td>\n<\/tr>\n<tr>\n<td>22. <a href=\"https:\/\/www.oreilly.com\/radar\/when-smaller-models-win\/\">When Smaller Models Win<\/a><\/td>\n<td class=\"src\">O&rsquo;Reilly Radar<\/td>\n<td class=\"dt\">Aug 26, 2026<\/td>\n<\/tr>\n<tr>\n<td>23. <a href=\"https:\/\/www.infoworld.com\/article\/4210689\/ai-inference-5-best-practices.html\">AI inference: 5 best practices<\/a><\/td>\n<td class=\"src\">InfoWorld<\/td>\n<td class=\"dt\">Aug 19, 2026<\/td>\n<\/tr>\n<\/table>\n<h4>Agent authority in practice<\/h4>\n<div class=\"cluster-intro\">The governance read of the week: what happens to delegated authority as an agent runs, and why the principal you granted access to is not always the principal acting.<\/div>\n<table class=\"index-table\">\n<tr>\n<th>Article<\/th>\n<th>Source<\/th>\n<th>Published<\/th>\n<\/tr>\n<tr>\n<td>24. <a href=\"https:\/\/www.oreilly.com\/radar\/principal-drift-in-practice\/\">Principal Drift in Practice<\/a><\/td>\n<td class=\"src\">O&rsquo;Reilly Radar<\/td>\n<td class=\"dt\">Aug 20, 2026<\/td>\n<\/tr>\n<\/table>\n<p>            <!-- Detailed write-ups --><\/p>\n<h2>Detailed write-ups<\/h2>\n<div class=\"article\">\n<h4>1. Elastic closes Deductive AI &mdash; observability buys the reasoning layer<\/h4>\n<p class=\"meta\">Business Wire &middot; Business Insider &middot; August 24&ndash;26, 2026<\/p>\n<p><strong>Elastic<\/strong> completed its acquisition of <strong>Deductive AI<\/strong> on <strong>August 24<\/strong>, on undisclosed terms, and what it bought is specific: an AI SRE agent that, in Elastic&rsquo;s description, &ldquo;gathers evidence, forms and tests hypotheses, and reasons across code, telemetry, and organizational knowledge to determine root cause,&rdquo; using reinforcement learning to improve its investigations over time. Elastic CEO <strong>Ash Kulkarni<\/strong> gave the rationale in a line &mdash; engineering teams are &ldquo;drowning in telemetry but starved for answers&rdquo; &mdash; and Deductive AI cofounder and former CEO <strong>Rakesh Kothari<\/strong> put the same point from the other side, that teams deserve better than hours spent manually tracing a root cause. The stated integration path matters more than either quote: the agent is to be paired with Elastic Observability&rsquo;s ability to infer entities, relationships and significant operational events from telemetry, which means the reasoning is meant to run over a topology Elastic already derives rather than over raw logs.<\/p>\n<p>The pattern to watch is that observability vendors have stopped competing on how much data they can hold. Storage and ingest are commoditising, the cost curve on both is under active downward pressure from customers, and no buyer in 2026 is choosing a platform because it can index more. The competition has moved to time-to-explanation &mdash; the distance between a page firing and an engineer understanding why. That is a reasoning problem, which is why acquisitions in this category now look like AI acquisitions rather than infrastructure ones. The reinforcement-learning detail is the part to press a vendor on at renewal: a system that learns from investigations is only as good as the corpus it learns from, so the useful question is whether the capability is grounded in <em>your<\/em> topology, change history and past incidents, or whether it is a generic model reading your logs with a confident tone.<\/p>\n<p><strong>Grafana Labs<\/strong> supplied the demand-side evidence in the same week, with CEO Raj Dutt telling Business Insider that annual recurring revenue has reached <strong>$600M<\/strong>, with AI a material driver of the growth. That direction of causation is worth stating plainly, because it is the commercial case for everything else in this issue: more AI in production means more services, more signals, more non-determinism and more people trying to understand a system they did not write. Observability spend is rising as a direct function of AI adoption. If your organisation is budgeting AI programmes without a corresponding line for the telemetry needed to operate them, the number is wrong, and this week&rsquo;s consolidation news is the market telling you so.<\/p>\n<p style=\"font-size:13px;color:#6b7280;margin:0;\">Sources: <a href=\"https:\/\/www.businesswire.com\/news\/home\/20260824863894\/en\/Elastic-Completes-Acquisition-of-Deductive-AI\">Business Wire (Elastic completes Deductive AI acquisition)<\/a> &middot; <a href=\"https:\/\/www.businessinsider.com\/grafana-labs-arr-600-million-ai-demand-2026-8\">Business Insider (Grafana ARR hits $600M)<\/a><\/p>\n<\/p><\/div>\n<div class=\"article\">\n<h4>2. Observability has a data problem, and telemetry pipelines are the answer on offer<\/h4>\n<p class=\"meta\">The New Stack &middot; August 25&ndash;26, 2026<\/p>\n<p>The New Stack&rsquo;s framing is the one to carry into your next capacity conversation: observability already has a data problem, and AI is about to make it worse. Will Kelly builds the case around <strong>Bronto<\/strong>, a Dublin data-observability startup whose co-CEOs <strong>Trevor Parsons<\/strong> and <strong>Noel Ruane<\/strong> argue that <strong>OpenTelemetry<\/strong> settled the collection question and left the storage question wide open &mdash; the company&rsquo;s head of community, <strong>Severin Neumann<\/strong>, is himself an OpenTelemetry maintainer. The numbers are the useful part: observability now consumes <strong>10&ndash;30% of infrastructure spend<\/strong> and still returns only partial access to the data, which is why every standard response is subtractive &mdash; sample the traces, index less of the logs, ratchet retention from 30 days to seven to three, and rehydrate from cold storage when an incident demands it. Bronto&rsquo;s own answer is a polymorphic datastore, <strong>BrontoDB<\/strong>, which it claims holds 100x more observability data than incumbents such as Datadog at twelve-month full-fidelity retention with sub-second search. Treat the multiple as a vendor claim; treat the diagnosis as sound, because agentic systems emit more steps and higher cardinality per unit of user-visible work than the request-response services they sit beside.<\/p>\n<p>The companion piece &mdash; Megan Carnegie&rsquo;s, sponsored by <strong>Apica<\/strong> &mdash; is the practical response, and its survey figures are the sharpest cost evidence of the week. <strong>54%<\/strong> of enterprises saw telemetry volume triple over the past year, <strong>43%<\/strong> of that growth came from AI\/ML workloads, average observability spending sits at <strong>$3.17 million<\/strong> and is growing <strong>28%<\/strong> year over year, and respondents expect a <strong>9.5x<\/strong> increase in telemetry data within two years, with <strong>44%<\/strong> bracing for somewhere between 6x and 100x. The finding that should change a roadmap conversation is this one: <strong>59%<\/strong> of organisations have already terminated or delayed an agentic AI deployment because of monitoring costs. That decision is being made in finance, not engineering. Apica CPTO <strong>Andi Mann<\/strong> makes the pipeline-first case, and the mechanics are vendor-neutral even where the framing is not &mdash; sample repetitive successful events while retaining every failure; enrich records with agent, session, model, tool, token and cost attributes so spend can be attributed at all; redact sensitive prompts and identifiers before they land; and aggregate metrics to destinations with different cost and retention profiles. The claimed payoff (40% lower total cost of ownership, and pipeline-mature organisations 80% more likely to avoid operational cost problems) is Apica&rsquo;s. The structural point is yours to keep: a pipeline is a vendor-independent place to make these decisions, which is also what makes an observability contract renegotiable.<\/p>\n<p>The third piece is the one most teams have not internalised yet, and <strong>Manveer Chawla<\/strong> &mdash; Zenith cofounder, previously a director of engineering at Confluent &mdash; states it plainly: for many agentic products the execution record &ldquo;turns out to be application data with a telemetry-shaped workload.&rdquo; The volumes are why. <strong>Laminar<\/strong> reports more than <strong>500,000 browser events per day<\/strong>, and a single browser-agent session running thirty minutes or more can generate hundreds of thousands of DOM diff events &mdash; a shape <strong>Postgres<\/strong> handles badly the moment users start asking for complete task histories and cross-run analysis, which is why the piece traces the migration towards an analytical store such as <strong>ClickHouse<\/strong>, with the <strong>OpenTelemetry GenAI semantic conventions<\/strong> and tools like <strong>Langfuse<\/strong> supplying the schema vocabulary. The consequence is not really a storage decision, it is a classification one: an agent trace records which tools were called, what context was retrieved and what was executed on a user&rsquo;s behalf, which makes it evidence in a dispute, an audit or a post-incident review of an automated action. That pulls retention, schema stability and access control out of ops and into legal and product. The cheap move this quarter is to decide deliberately which spans are ephemeral telemetry and which are records, and route them differently in the pipeline from the start.<\/p>\n<p style=\"font-size:13px;color:#6b7280;margin:0;\">Sources: <a href=\"https:\/\/thenewstack.io\/opentelemetry-observability-telemetry-storage\/\">The New Stack (observability has a data problem)<\/a> &middot; <a href=\"https:\/\/thenewstack.io\/agentic-ai-telemetry-costs\/\">The New Stack (telemetry pipelines and agent costs)<\/a> &middot; <a href=\"https:\/\/thenewstack.io\/agent-traces-application-data\/\">The New Stack (when agent traces become application data)<\/a><\/p>\n<\/p><\/div>\n<div class=\"article\">\n<h4>3. The AI line item goes variable, and nobody can prove the return yet<\/h4>\n<p class=\"meta\">DevOps.com &middot; InfoWorld &middot; SiliconANGLE &middot; August 24&ndash;26, 2026<\/p>\n<p><strong>DevOps.com<\/strong>&rsquo;s Carla Castillo named the structural change cleanly: AI coding spend now behaves more like cloud infrastructure than SaaS. The concrete trigger is <strong>GitHub Copilot<\/strong>&rsquo;s move to usage-based billing in mid-2026, pricing premium usage against a metered pool of credits that premium models and agentic features draw down &mdash; which means, as the piece puts it, that two engineers on the same plan can generate very different costs, with surprises measured in tens of dollars per user per month. Engineering leaders who spent a decade learning to forecast tooling spend are now holding a line item that behaves like egress. Castillo&rsquo;s four metrics are the minimum instrumentation: cost per developer across all tools, utilisation (active versus assigned seats), premium-model usage share, and forecast versus budget. If your finance partner is still carrying AI tooling as a fixed cost, that is a conversation to have before the next quarter closes rather than after.<\/p>\n<p><strong>InfoWorld<\/strong> attacked the same total from the other side, and its answer is more specific than the headline suggests. Joseph Morais argues the biggest hidden AI cost is not GPUs but <strong>context size and data quality<\/strong>: &ldquo;if you&rsquo;re not preparing and slimming down the data that feeds into a model&rsquo;s context, you&rsquo;re merely paying to process noise.&rdquo; With <strong>73%<\/strong> of enterprises reporting that their AI costs have already outpaced what they budgeted, that is the cheapest lever most teams have not pulled, because it sits upstream of everything else in this issue &mdash; a bloated context inflates token spend, inference latency and the telemetry emitted per call simultaneously. The prescription is unglamorous and familiar to anyone who has run a streaming platform: filter and shape data before it reaches the model, using stream processing (<strong>Apache Flink<\/strong> is the worked example), and enforce data contracts through schema validation so the filtering does not silently rot. The author writes from <strong>Confluent<\/strong>, so the tooling choice is not neutral &mdash; but the underlying correction holds. A cost model denominated in GPU hours alone will underestimate, and it will point optimisation effort at the one component you probably cannot change.<\/p>\n<p>Which sets up the week&rsquo;s hardest question, also from InfoWorld: enterprises can measure AI usage, but proving it delivered value remains largely unsolved. Taryn Plumb stacks the evidence &mdash; <strong>IBM<\/strong> finds only <strong>29%<\/strong> of executives can confidently measure AI ROI and just <strong>25%<\/strong> of AI initiatives deliver the return expected, while <strong>Kyndryl<\/strong> reports <strong>61%<\/strong> of senior business leaders feel more burdened to prove AI ROI than they did a year ago. <strong>Tempo<\/strong> CEO <strong>Vic Chynoweth<\/strong> says the quiet part out loud: &ldquo;The amount of money people are spending on AI is enormous, and a very large percentage of it is wasted.&rdquo; His company&rsquo;s answer, launched this week as <strong>Workforce Intelligence<\/strong> on the Atlassian Marketplace, correlates inference API telemetry from tools such as <strong>GitHub Copilot<\/strong>, <strong>OpenAI Codex<\/strong> and <strong>Claude Code<\/strong> with human effort data, attaching cost and performance figures to individual <strong>Jira<\/strong> issues; CTO <strong>Shams Chauthani<\/strong>&rsquo;s claim is that it &ldquo;correlates the session to the commit directly, so the answer is verifiable.&rdquo; Tempo has tracked human-delivered work in Jira since 2007 and counts 30,000-plus customers including Cisco, Airbus and Oracle, which is the real asset: the denominator already exists. The join is imperfect &mdash; ticket-level attribution inherits every flaw in your ticket hygiene and misses AI-assisted work that never became an issue &mdash; but the pressure behind it is real, with <strong>Forrester<\/strong> finding fewer than a third of decision-makers can tie AI value to financial growth and expecting enterprises to defer a quarter of planned AI spending into 2027. If you take one action from this cluster, make it this: pick a unit of delivered work your organisation already counts, and start attributing AI spend to it now, imperfectly, rather than waiting for a measurement framework that will not arrive.<\/p>\n<p style=\"font-size:13px;color:#6b7280;margin:0;\">Sources: <a href=\"https:\/\/devops.com\/your-ai-coding-budget-is-becoming-a-variable-cloud-bill\/\">DevOps.com (AI coding budget as a variable cloud bill)<\/a> &middot; <a href=\"https:\/\/www.infoworld.com\/article\/4210670\/why-your-biggest-hidden-ai-cost-isnt-gpus.html\">InfoWorld (your biggest hidden AI cost isn&rsquo;t GPUs)<\/a> &middot; <a href=\"https:\/\/www.infoworld.com\/article\/4213146\/enterprises-can-measure-ai-usage-but-the-hard-part-is-proving-that-it-actually-delivered-value.html\">InfoWorld (measuring AI usage vs proving value)<\/a> &middot; <a href=\"https:\/\/siliconangle.com\/2026\/08\/25\/tempo-launches-workforce-intelligence-to-tie-ai-spend-to-jira-work-items\/\">SiliconANGLE (Tempo Workforce Intelligence)<\/a><\/p>\n<\/p><\/div>\n<div class=\"article\">\n<h4>4. Serving the load: shadow-engine recovery, dynamic capacity, and the limits of real time<\/h4>\n<p class=\"meta\">NVIDIA Technical Blog &middot; Google Cloud Blog &middot; The New Stack &middot; InfoWorld &middot; August 23&ndash;26, 2026<\/p>\n<p><strong>NVIDIA<\/strong> published the most operationally specific piece of the week, and the headline number is the whole argument: in its own benchmark, restoring a failed inference worker took <strong>283 seconds<\/strong> by cold restart and <strong>7.3 seconds<\/strong> with shadow engine recovery in <strong>Dynamo<\/strong> &mdash; roughly a 39x improvement. The mechanism explains why inference recovery is slow to begin with. A shadow engine sits fully initialised but idle on the <em>same<\/em> GPUs as the active engine, and a per-GPU sidecar called the <strong>GPU Memory Service<\/strong>, built on the CUDA Virtual Memory Management API, lets both engines share a single copy of the weights instead of duplicating them in HBM; engine election is arbitrated with a POSIX <code>flock<\/code>, and NCCL communicators and captured CUDA graphs are preserved rather than rebuilt. Recovery becomes a handoff instead of a rebuild, and the tail behaviour shows it: across the post-fault window, p50 time-to-first-token was <strong>1,311 ms<\/strong>, decode held at <strong>46 tokens\/s\/user<\/strong>, and just <strong>1 of 398 requests<\/strong> exceeded a five-second TTFT. The rig was two workers serving <strong>GLM-5.2<\/strong> quantised to <strong>NVFP4<\/strong> on <strong>NVIDIA B200<\/strong> nodes at TP=8 with a 200K maximum context and FP8 KV cache, driven at 0.7 requests per second with 32,000 input and 1,000 output tokens per request. Two constraints before you plan around it: the feature is a preview, and it requires <strong>Kubernetes 1.34 or newer<\/strong> with DRA enabled and the NVIDIA GPU DRA driver installed, with vLLM as the primary supported backend. The lesson generalises past Dynamo: capacity planning for inference belongs in units of <em>time to restore serving capacity<\/em>, not replica count, and if you have never measured that number on your own stack it is the single most useful experiment on this list.<\/p>\n<p><strong>Google Cloud<\/strong>&rsquo;s <strong>Drew Bradstock<\/strong>, senior director of product for orchestration and Kubernetes, supplied the planning-side companion and opened with the gap that motivates it: <strong>90%<\/strong> of enterprises want to deploy agents within the next three years, while only <strong>17%<\/strong> of IT leaders are confident their current IT setup can handle the load. The post is essentially a catalogue of consumption modes and when each applies &mdash; <strong>Dynamic Workload Scheduler<\/strong> in <strong>calendar mode<\/strong> for capacity you can book against a known date, <strong>flex-start<\/strong> for jobs that can wait for a window, Spot VMs for work that can be displaced, and committed use discounts for the predictable floor, with compute flexible CUDs quoted at up to <strong>63%<\/strong> off. On the scheduling side it points at GKE <strong>custom ComputeClasses<\/strong> and dynamic resource allocation so that fallback priority is expressed in the cluster rather than in a runbook, with managed instance groups, instance flexibility and bulk VM creation underneath. The useful reframe is treating accelerator capacity as a portfolio with an explicit displacement policy rather than as a single autoscaling decision &mdash; anyone who has watched a batch evaluation starve an interactive endpoint at the wrong moment will recognise the failure mode it is trying to prevent.<\/p>\n<p>The other two pieces set the boundaries. <strong>ScyllaDB<\/strong>&rsquo;s Felipe Cardeneti Mendes asks why real-time AI at scale is so hard and answers with measurements rather than adjectives: a system holding p99 under <strong>10 ms<\/strong> at a concurrency of 300 degraded to a <strong>three-second p99<\/strong> at around <strong>740,000 operations per second<\/strong>, and vector-index mutations under load dropped recall to <strong>42%<\/strong> while serving roughly <strong>150,000 approximate-nearest-neighbour queries per second<\/strong>. The model did not change; the answers got worse anyway. His line is the one to quote in design review &mdash; &ldquo;Tail latency isn&rsquo;t a bug that you can fix, it&rsquo;s a property of your architecture&rdquo; &mdash; and the prescriptions follow from it: isolate training from serving, index vectors asynchronously, and monitor feature freshness and index health as first-class signals rather than afterthoughts. Adding accelerators does not fix a variance problem. <strong>InfoWorld<\/strong>&rsquo;s Bill Doerrfeld frames the market around all of it: <strong>neocloud<\/strong> revenue passed <strong>$25 billion<\/strong> in 2025, and Gartner senior principal analyst <strong>Hardeep Singh<\/strong> expects the specialists &mdash; CoreWeave, Lambda, Nebius, RunPod, Vultr &mdash; to take about <strong>20%<\/strong> of a <strong>$267 billion<\/strong> AI cloud market by 2030. The concentration risk is the part a platform team should read twice: <strong>Microsoft<\/strong> accounted for <strong>67%<\/strong> of CoreWeave&rsquo;s 2025 revenue, and CoreWeave is committing roughly <strong>$30 billion<\/strong> of capex in 2026 against that customer base. The article lands on coexistence &mdash; neoclouds taking frontier training, hyperscalers keeping the broader enterprise estate &mdash; which for you means optionality: if your inference stack only runs on one provider&rsquo;s scheduler, you have no leverage whichever way it settles.<\/p>\n<p style=\"font-size:13px;color:#6b7280;margin:0;\">Sources: <a href=\"https:\/\/developer.nvidia.com\/blog\/restore-llm-inference-capacity-in-seconds-with-shadow-engine-recovery-in-nvidia-dynamo\">NVIDIA (shadow engine recovery in Dynamo)<\/a> &middot; <a href=\"https:\/\/cloud.google.com\/blog\/topics\/ai-infrastructure\/best-practices-for-dynamic-capacity-management\/\">Google Cloud (dynamic capacity management)<\/a> &middot; <a href=\"https:\/\/thenewstack.io\/real-time-ai-scale\/\">The New Stack (why real-time AI at scale is so hard)<\/a> &middot; <a href=\"https:\/\/www.infoworld.com\/article\/4211105\/can-neoclouds-corner-ai-compute.html\">InfoWorld (can neoclouds corner AI compute?)<\/a><\/p>\n<\/p><\/div>\n<div class=\"article\">\n<h4>5. Where agents actually run: the missing runtime, the sandbox, and the platform&rsquo;s job<\/h4>\n<p class=\"meta\">DevOps.com &middot; Google Cloud Blog &middot; The New Stack &middot; Blocks &amp; Files &middot; August 25&ndash;29, 2026<\/p>\n<p><strong>DevOps.com<\/strong>&rsquo;s Anuj Kapoor made the case that long-running AI agents have no proper runtime, and the framing is sharper than &ldquo;agents need state&rdquo;. The execution models we have assume either a request that finishes in seconds holding nothing, or a batch job that runs to completion and dies; an agent working a task across hours, external dependencies and human decisions is neither, and plain HTTP request\/response is the specific abstraction he names as the wrong shape. His list of what such a runtime must provide is a serviceable checklist to run your own setup against: &ldquo;state, retries, checkpoints and recovery&rdquo;, the ability to &ldquo;coordinate multi-step workflows, survive failures, pause for human review and resume reliably&rdquo;, and around all of it observability, correlation identifiers, retry policies, cost controls, security boundaries and careful versioning. His proposed answer is durable workflow orchestration, worked through <strong>Azure Durable Functions<\/strong> and <strong>Azure AI Foundry<\/strong> &mdash; the pattern generalises to any durable-execution engine, and the discipline worth adopting is to write down which of those properties your current assembly of queues, workflow engines and databases actually delivers today.<\/p>\n<p><strong>Google Cloud<\/strong> and <strong>Anyscale<\/strong> published the security half of the same problem jointly &mdash; Google staff software engineer <strong>Andrew Sy Kim<\/strong> with Anyscale CTO <strong>Philipp Moritz<\/strong> &mdash; and the number they lead with is the reason to pay attention: Ray on GKE scaled to <strong>100,000 gVisor sandboxes in 17.3 seconds<\/strong> across thousands of nodes. The API arrived in <strong>Ray 2.58<\/strong> as <code>ray.experimental.sandbox<\/code>: create an environment from an OCI image with explicit CPU and memory limits, execute commands inside it, move files in and out, inspect its state, tear it down. <strong>gVisor<\/strong>&rsquo;s user-space kernel gives a materially smaller attack surface than container isolation alone, with sub-second sandbox startup and low per-sandbox memory overhead, and critically it does not require exposing a Docker daemon or the host Docker socket &mdash; which is precisely how most home-grown code-execution sandboxes end up leaking. Ray is where a great deal of distributed AI work already runs, and increasingly what runs on it is code a model wrote; Kata Containers support is signposted as future work. If you are executing agent-generated or customer-supplied code anywhere near your data, this is the shape of the answer.<\/p>\n<p>The New Stack contributed the platform-side piece. <strong>Port<\/strong>&rsquo;s Matar Peles sets out three roles agents play in a developer platform &mdash; &ldquo;AI agents as platform consumers&rdquo;, &ldquo;AI agents as internal platform components&rdquo;, and &ldquo;AI as a resource with its own lifecycle&rdquo;, which she labels AgenticOps &mdash; a genuinely useful taxonomy for arguing about scope, because the governance an agent needs when it reads context differs sharply from what it needs when it is provisioned as a managed resource; she notes <strong>47%<\/strong> of organisations surveyed in early 2026 asked for agent and skill registries, which is the registry problem arriving on schedule. <strong>Nutanix<\/strong>, meanwhile, shipped the enterprise-infrastructure version of the same bet: <strong>Nutanix Enterprise AI 2.8<\/strong> and <strong>Service Provider Central<\/strong> generally available, <strong>Nutanix Kubernetes Platform 2.19<\/strong> arriving shortly, plus <strong>Nutanix Agent Gateway<\/strong> and <strong>Nutanix Private Inference<\/strong>, with NVIDIA AI Enterprise integration and Kubeflow, Milvus and Slurm in the applications catalogue; speculative decoding is claimed to accelerate token generation up to <strong>2.5x<\/strong>, and parameter-efficient fine-tuning is supported for models under <strong>8B parameters<\/strong>. EVP of product management <strong>Thomas Cornely<\/strong> put the strategy in one sentence: &ldquo;Enterprise AI should not require customers to rebuild the systems that already run their business.&rdquo; The through-line across all four is that the platform layer is being redefined around agents this year, and the teams that write down what their platform guarantees an agent will get &mdash; state, identity, isolation, observability &mdash; will have a much easier time than the ones discovering the gaps one incident at a time.<\/p>\n<p style=\"font-size:13px;color:#6b7280;margin:0;\">Sources: <a href=\"https:\/\/devops.com\/the-missing-runtime-for-long-running-ai-agents\/\">DevOps.com (the missing runtime for long-running AI agents)<\/a> &middot; <a href=\"https:\/\/cloud.google.com\/blog\/products\/containers-kubernetes\/gvisor-sandboxes-for-ray-clusters-on-gke\">Google Cloud (gVisor sandboxes for distributed Ray clusters)<\/a> &middot; <a href=\"https:\/\/thenewstack.io\/ai-agent-platform-roles\/\">The New Stack (the 3 roles AI agents play in your developer platform)<\/a> &middot; <a href=\"https:\/\/www.blocksandfiles.com\/hci\/2026\/08\/26\/nutanix-adds-more-rooms-to-its-agentic-ai-building\/5292580\">Blocks &amp; Files (Nutanix agentic AI)<\/a><\/p>\n<\/p><\/div>\n<div class=\"article\">\n<h4>6. Automated diagnosis is not automated understanding<\/h4>\n<p class=\"meta\">DevOps.com &middot; Dynatrace &middot; The Register &middot; August 25&ndash;26, 2026<\/p>\n<p>The best-titled piece of the week is also the most important for anyone shipping incident AI, and Jyostna Seelam&rsquo;s central assertion fits in four words: &ldquo;Correlation is not diagnosis.&rdquo; Her argument is that most AI incident tools group related alerts and present the grouping as a cause &mdash; &ldquo;grouping related symptoms isn&rsquo;t the same thing as identifying a cause&rdquo; &mdash; because &ldquo;finding symptoms is easy; understanding causes is not&rdquo;. Postmortem practice has spent two decades learning that the first plausible cause is usually not the whole cause, and that a confident wrong answer costs more than no answer at all: it is trusted early, is wrong once expensively, and is then ignored permanently. Four requirements fall out of the piece and together they make a serviceable vendor scorecard &mdash; test causal <em>direction<\/em> against a dependency graph rather than inferring it from co-occurrence; match the incident against incident history; communicate uncertainty instead of guessing confidently; and work from live dependency data rather than an architecture diagram that stopped being true two quarters ago. Her single best evaluation question is one to put to every root-cause feature arriving through acquisition, Elastic&rsquo;s included: &ldquo;Does this tool group related alerts, or does it actually explain why one caused another?&rdquo;<\/p>\n<p><strong>Dynatrace<\/strong>&rsquo;s State of SRE and Platform Engineering 2026 supplies the organisational context, and the sample is solid enough to argue with: <strong>919 senior IT leaders<\/strong> at enterprises above <strong>$500M in annual revenue<\/strong> across the Americas, EMEA and Asia-Pacific, surveyed between October 2025 and January 2026. The headline shift is that <strong>67%<\/strong> of SREs now name AI model monitoring their top use case and <strong>50%<\/strong> already use AI-powered capabilities for automated incident response &mdash; SRE has become an AI-operations function faster than most job descriptions have caught up. The strain shows in the supporting numbers: nearly half of SRE respondents say too many data sources and metrics hinder their ability to define and manage effective SLOs, only <strong>40%<\/strong> of platform engineers embed observability across all deployment stages, and <strong>37%<\/strong> name integrating with existing tools their top challenge &mdash; this despite <strong>89%<\/strong> of organisations practising platform engineering having implemented an internal developer platform and <strong>92%<\/strong> reporting executive support. Chief product officer <strong>Steve Tack<\/strong> frames the required response as a move &ldquo;from managing systems to orchestrating them, connecting observability, automation, and agentic AI&rdquo;. Read it against your own team, because the failure most organisations are heading for is a staffing and ownership one rather than a technology gap: AI services get shipped by teams that do not run them, and the operational load lands somewhere unplanned.<\/p>\n<p>The week&rsquo;s reminder that the substrate still needs tending came from <strong>The Register<\/strong>: <strong>Kubernetes 1.37<\/strong>, nicknamed <strong>Garhwal<\/strong>, landed on <strong>August 26<\/strong> with 67 changes and a clear-out of three pieces of legacy. The timelines matter more than the headline. <strong>kube-dns<\/strong> is being retired in favour of CoreDNS &mdash; the default since v1.13 &mdash; and anything still on it must move before <strong>1.40<\/strong>. The <strong>IPVS<\/strong> mode of kube-proxy is deprecated with removal signposted from <strong>1.43<\/strong> and nftables as the replacement, so anyone who tuned kube-proxy for scale years ago and never revisited it now has a migration to schedule. <strong>cgroup v1<\/strong> bites hardest, because it is a node-operating-system problem rather than a cluster one: since <strong>1.35<\/strong>, nodes relying on it will not initialise at all, which makes it a fleet-wide OS upgrade rather than a cluster-level change. On the additive side, the <strong>metrics.k8s.io<\/strong> Metrics API finally graduated to GA after nine years in beta, with release lead <strong>Dipesh Rawat<\/strong> noting it had been &ldquo;widely used in production for years&rdquo; regardless of the label. The practical step is small and worth doing now: audit your clusters for all three before you plan the upgrade, not during it. The clusters most likely to be affected are the oldest and least-touched &mdash; which, in a year when AI workloads are being scheduled onto whatever capacity exists, are increasingly the ones running something that matters.<\/p>\n<p style=\"font-size:13px;color:#6b7280;margin:0;\">Sources: <a href=\"https:\/\/devops.com\/automated-diagnosis-isnt-automated-understanding-what-postmortems-teach-us-about-building-trustworthy-incident-ai\/\">DevOps.com (automated diagnosis isn&rsquo;t automated understanding)<\/a> &middot; <a href=\"https:\/\/www.dynatrace.com\/news\/press-release\/state-of-sre-platform-engineering-2026\/\">Dynatrace (State of SRE and Platform Engineering 2026)<\/a> &middot; <a href=\"https:\/\/www.theregister.com\/devops\/2026\/08\/26\/kubernetes-cleans-house-bins-legacy-kube-dns-ipvs-and-cgroup-v1\/5292717\">The Register (Kubernetes 1.37 removes kube-dns, IPVS and cgroup v1)<\/a><\/p>\n<\/p><\/div>\n<div class=\"article\">\n<h4>7. Foundational: what a token really costs, when a smaller model wins, and the comprehension you traded away<\/h4>\n<p class=\"meta\">O&rsquo;Reilly Radar &middot; InfoWorld &middot; August 19&ndash;28, 2026<\/p>\n<p><strong>Tokens Aren&rsquo;t Dollars<\/strong>, by Tim O&rsquo;Brien, is the reference read behind every cost story in this issue, and its argument is narrower and more useful than the title suggests: counting tokens is necessary, but a token is not a unit of value, and tokens are not interchangeable across models or use cases. The same nominal volume can be cheap and productive on one model and an expensive mistake on another, which makes cross-model comparison by token count actively misleading &mdash; and none of it captures the retrieval, orchestration and human-review work wrapped around the call. His worked cases are all about mismatch rather than volume: spending hundreds of dollars a day on a task with no way to judge whether that is waste, or running a billion tokens through a frontier model for work a cheaper one would have finished. Read it alongside the Tempo launch and the pattern is consistent &mdash; the industry is groping towards a denominator, and per-token accounting is not it.<\/p>\n<p><strong>When Smaller Models Win<\/strong>, by Sruly Rosenblat, is the architectural counterpart and the highest-leverage lever most teams have not pulled. The examples are concrete rather than rhetorical: a <strong>4B-parameter<\/strong> research model, LiteResearcher, is reported to have beaten Claude Sonnet 4.5 on some search benchmarks; the <strong>Docling<\/strong> family does document extraction at as little as <strong>258 million parameters<\/strong>; and a 4B negotiation model outperformed frontier GPT-5-series models on social negotiation tasks. The author&rsquo;s own <strong>15M-parameter<\/strong> chess model predicted a human&rsquo;s next move with 27% accuracy &mdash; a reminder that task-shaped models buy cheaply what general-purpose ones spend heavily to approximate, with LoRA making the fine-tuning cheap enough to be worth an experiment. The discipline implied is to profile by <em>step<\/em> rather than by application: classification, extraction, routing and the intermediate hops of an agent chain rarely need frontier reasoning, and they usually arrive there by default rather than by decision. <strong>InfoWorld<\/strong>&rsquo;s Isaac Sacolick covers the serving side with five practices &mdash; architect for integration and performance, secure the AI&rsquo;s data and actions, separate training and inference requirements, design for flexible and resilient operations, and optimise for costs and changing AI models &mdash; and lands two numbers worth carrying into planning: only <strong>25%<\/strong> of organisations have moved 40% or more of their AI experiments into production, and a multi-agent system can consume <strong>15 times as many tokens<\/strong> as a single chat interaction. That multiplier is the argument for step-level model selection restated as arithmetic.<\/p>\n<p>The fourth read is the governance one, and it is not the identity paper its title suggests. <strong>Principal Drift in Practice<\/strong>, by Shreshta Shyamsundar, is about <em>cognitive debt<\/em> &mdash; the widening gap between how complex a system has become and how well the team that owns it still understands it &mdash; and the loss of control that follows when an organisation ships code it can no longer reason about. Her claim is that the 2024&ndash;2025 preference for speed over comprehension has surfaced in 2026 as tripled production incident rates and as architectural decisions nobody on the team is equipped to make. The mitigations are process rather than tooling, and they are implementable next sprint: route code review by tier, reserving full line-by-line review for security, money movement and data integrity while decoupled changes and utilities get a systems-level inspection; never let the agent that authored a change be its sole reviewer; and use literate code explanations with comprehension checkpoints so understanding is demonstrated rather than assumed. To make it stick she wants executive sponsorship, a written tier-assessment policy, CI\/CD enforcement, and postmortems that check whether the tier was assigned correctly in the first place. For anyone operating AI in production this is the read that connects to everything else in the issue: an incident AI you cannot second-guess, an agent runtime you did not design and a telemetry pipeline you inherited are the same debt wearing different clothes.<\/p>\n<p style=\"font-size:13px;color:#6b7280;margin:0;\">Sources: <a href=\"https:\/\/www.oreilly.com\/radar\/tokens-arent-dollars\/\">O&rsquo;Reilly Radar (Tokens Aren&rsquo;t Dollars)<\/a> &middot; <a href=\"https:\/\/www.oreilly.com\/radar\/when-smaller-models-win\/\">O&rsquo;Reilly Radar (When Smaller Models Win)<\/a> &middot; <a href=\"https:\/\/www.infoworld.com\/article\/4210689\/ai-inference-5-best-practices.html\">InfoWorld (AI inference: 5 best practices)<\/a> &middot; <a href=\"https:\/\/www.oreilly.com\/radar\/principal-drift-in-practice\/\">O&rsquo;Reilly Radar (Principal Drift in Practice)<\/a><\/p>\n<\/p><\/div>\n<p>            <!-- Calls to action \/ Watch list --><\/p>\n<div class=\"watchlist\">\n<h2>Calls to action<\/h2>\n<ul>\n<li><strong>Put your AI tooling spend on the same footing as cloud spend.<\/strong> It is now a variable bill, not a per-seat licence. Give it per-team attribution, a budget with an alert threshold, a named owner and a monthly review &mdash; and do it before the next quarter closes rather than after the first spike.<\/li>\n<li><strong>Build the cost model on more than GPU hours.<\/strong> InfoWorld&rsquo;s answer to where the money actually goes is context size and data quality &mdash; if you are not preparing and slimming the data that feeds the context window, you are paying to process noise. Audit your prompt and retrieval payload sizes before you go shopping for cheaper compute; a GPU-denominated model will underestimate and will point your optimisation effort at the wrong thing.<\/li>\n<li><strong>Pick a denominator for AI value this quarter, imperfectly.<\/strong> Usage telemetry is abundant and is not value. Attribute spend to a unit of delivered work your organisation already counts &mdash; tickets, merged changes, closed incidents &mdash; and accept that version one will be wrong. It is still more defensible than a productivity percentage with nothing underneath it.<\/li>\n<li><strong>Treat the telemetry pipeline as a cost-control tier.<\/strong> Tail-based sampling, edge aggregation, routing audit-grade data to object storage and only the hot subset to the expensive index, and redaction of agent inputs and outputs. It also gives you a vendor-independent place to make those decisions, which is what makes an observability renewal negotiable.<\/li>\n<li><strong>Classify agent traces as records, not just telemetry.<\/strong> Decide now which spans are ephemeral debugging data and which are the evidence of what an automated system decided on a customer&rsquo;s behalf &mdash; then give the second group a stable schema, a legally-set retention period and access controls to match, and route them separately.<\/li>\n<li><strong>Measure your time to restore inference capacity.<\/strong> An inference replica is not a stateless web replica; the expensive part is loaded weights and warm cache. NVIDIA&rsquo;s shadow-engine work exists because that recovery is slow. Run the experiment on your own stack and put the number in your capacity plan.<\/li>\n<li><strong>Audit for kube-dns, IPVS proxy mode and cgroup v1 before planning the 1.37 upgrade.<\/strong> All three are gone in Kubernetes 1.37, and cgroup v1 removal is a node-OS problem rather than a cluster one. The clusters most likely to be affected are the oldest and least-touched &mdash; which is increasingly where spare capacity for AI workloads is being found.<\/li>\n<li><strong>Profile your workload by step and move the defaults.<\/strong> Classification, extraction, routing and intermediate agent steps rarely need a frontier model. Measure which calls genuinely require frontier reasoning rather than assuming the application-level answer, then change the routing defaults for everything else.<\/li>\n<\/ul><\/div>\n<div class=\"watchlist\">\n<h2>On our watch list<\/h2>\n<ul>\n<li><strong>Whether observability consolidation delivers reasoning or just packaging.<\/strong> Elastic has bought Deductive AI&rsquo;s root-cause capability. Watch whether the integrated product can ground its explanations in customer-specific topology and change history, or whether it degrades into a generic model reading logs with a confident tone &mdash; the difference determines whether the whole acquisition wave was worth it.<\/li>\n<li><strong>Observability spend as a fixed percentage of AI spend.<\/strong> Grafana&rsquo;s $600M ARR on AI-driven demand suggests a ratio is forming. Watch whether organisations start budgeting telemetry as an explicit fraction of every AI programme, or keep discovering the cost after the workload is already in production.<\/li>\n<li><strong>A real runtime for long-running agents.<\/strong> Durable state, pause-and-resume without holding compute, checkpointed recovery and persistent identity across a multi-hour run. Everyone is currently assembling this from workflow engines and queues. Watch for the first credible product category, and for whether the hyperscalers or the workflow vendors get there first.<\/li>\n<li><strong>Attribution models for AI value.<\/strong> Tempo&rsquo;s Jira join is one attempt; it will not be the only one. Watch which unit of work the market settles on &mdash; tickets, merged changes, resolved incidents, customer outcomes &mdash; because whatever wins will shape how AI programmes are funded for the next several years.<\/li>\n<li><strong>Calibrated uncertainty in incident AI.<\/strong> The postmortem argument is that a confident wrong diagnosis is worse than none. Watch whether vendors start shipping confidence signals and evidence trails alongside root-cause conclusions, or whether the first wave of incident AI burns its credibility the way early anomaly detection did.<\/li>\n<li><strong>Sandboxing becoming the default for agent-generated code.<\/strong> gVisor on distributed Ray clusters is a strong pattern. Watch whether user-space kernel isolation becomes the expected posture for running model-written code, or whether teams keep relying on container boundaries that were never designed for unreviewed code.<\/li>\n<li><strong>Cognitive debt showing up in a named incident.<\/strong> The argument that teams are shipping systems they no longer understand &mdash; and that this is behind tripled production incident rates &mdash; is still an argument rather than a case file. Watch for the first well-documented postmortem that names comprehension rather than tooling as the root cause, because that is what will move tier-based review and builder\/reviewer separation from a good idea to a control someone has to sign off on.<\/li>\n<li><strong>Whether neoclouds hold their position as inference outgrows training.<\/strong> The specialist bet is that AI-native economics beat general-purpose cloud; the hyperscaler bet is data gravity and integration. Watch pricing, capacity availability and whether serious workloads start running multi-provider by design rather than by accident.<\/li>\n<\/ul><\/div>\n<\/td>\n<\/tr>\n<p>        <!-- Footer --><\/p>\n<tr>\n<td class=\"footer\">\n<p class=\"brand\">AI Ops<\/p>\n<p>A weekly intelligence bulletin from Security Radar LLC.<br \/>\n            Curated by Paul Davis &middot; <a href=\"mailto:paul.davis@security-radar.com\">paul.davis@security-radar.com<\/a><\/p>\n<p>&copy; 2026 Security Radar LLC. All rights reserved.<\/p>\n<p>Article titles and summaries are excerpted for review and commentary; all linked articles remain the copyright of their respective publishers and authors.<\/p>\n<p>*|LIST:ADDRESS|*<\/p>\n<p><a href=\"*|ARCHIVE|*\">View this email in your browser<\/a> &middot; <a href=\"*|UNSUB|*\">Unsubscribe<\/a><\/p>\n<\/td>\n<\/tr>\n<\/table>\n<\/td>\n<\/tr>\n<\/table>\n","protected":false},"excerpt":{"rendered":"<p>August 30, 2026 &middot; Weekly Edition AI Ops This week the beat narrowed to a single question: what does it actually cost to run AI in production, and can you see it? Elastic closed its acquisition of Deductive AI and bought itself an incident-reasoning layer; Grafana said AI demand pushed&#8230;<\/p>\n","protected":false},"author":1,"featured_media":0,"comment_status":"closed","ping_status":"","sticky":false,"template":"","format":"standard","meta":{"footnotes":""},"categories":[48],"tags":[],"class_list":["post-5781","post","type-post","status-publish","format-standard","hentry","category-ai-ops"],"_links":{"self":[{"href":"https:\/\/www.cybersecurityinstitute.com\/blog\/index.php?rest_route=\/wp\/v2\/posts\/5781","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/www.cybersecurityinstitute.com\/blog\/index.php?rest_route=\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/www.cybersecurityinstitute.com\/blog\/index.php?rest_route=\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/www.cybersecurityinstitute.com\/blog\/index.php?rest_route=\/wp\/v2\/users\/1"}],"replies":[{"embeddable":true,"href":"https:\/\/www.cybersecurityinstitute.com\/blog\/index.php?rest_route=%2Fwp%2Fv2%2Fcomments&post=5781"}],"version-history":[{"count":1,"href":"https:\/\/www.cybersecurityinstitute.com\/blog\/index.php?rest_route=\/wp\/v2\/posts\/5781\/revisions"}],"predecessor-version":[{"id":5812,"href":"https:\/\/www.cybersecurityinstitute.com\/blog\/index.php?rest_route=\/wp\/v2\/posts\/5781\/revisions\/5812"}],"wp:attachment":[{"href":"https:\/\/www.cybersecurityinstitute.com\/blog\/index.php?rest_route=%2Fwp%2Fv2%2Fmedia&parent=5781"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/www.cybersecurityinstitute.com\/blog\/index.php?rest_route=%2Fwp%2Fv2%2Fcategories&post=5781"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/www.cybersecurityinstitute.com\/blog\/index.php?rest_route=%2Fwp%2Fv2%2Ftags&post=5781"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}