8news

Tech • AI • Robotics

VIDEO
ENFR
TodayShortsTop StoriesFor youTopicsVideosYT channelsArchivesSearchFavorites

Daily Podcast full article

Cisco says agents dominate inference as AI’s workload moves from chat to orchestration

Cisco’s claim that AI agents now account for roughly 60% of global inference reframes the industry’s bottleneck: not the one-off chatbot prompt, but persistent, tool-using software that loops, calls models repeatedly and forces enterprises to rethink networking, identity, observability and cost control.

Generated September 10, 2026 at 2:36 AM UTC1418 words
AI-generated illustration

The headline Cisco wanted investors to hear

Cisco’s latest AI argument is blunt: agents, not chat prompts, have become the dominant consumer of inference. At the Goldman Sachs Communacopia + Technology Conference on September 8, Cisco President and Chief Product Officer Jeetu Patel said agent token consumption has risen 14-fold since February and now represents about 60% of total inference capacity or token volume . The company’s investor page confirms that Patel appeared alongside Chief Financial Officer Mark Patterson at the conference, giving the remarks the status of an investor-facing thesis rather than a passing product comment .

That distinction matters. Cisco is not merely saying more people are using AI assistants. It is saying the character of AI demand has changed. A chatbot session is often bursty: a user asks, waits, edits and stops. An agentic workflow can run across many steps, call tools, retrieve context, inspect results, retry a failed action and hand off to another system. Each loop creates more tokens, more model calls and more traffic across corporate and cloud networks.

Patel tied the 60% agent figure to OpenRouter data and said agents had already overtaken humans in token consumption earlier this year . GuruFocus, in a September 9 report focused on Cisco’s comments, likewise described AI agents as consuming roughly 60% of global inference capacity and noted the same 14-fold surge in token use since February . The point is less that one measurement captures the whole market perfectly than that Cisco now treats machine-to-machine AI work as the leading demand signal.

From model training glamour to inference economics

The AI industry has spent years narrating progress through training: bigger clusters, larger models, more expensive runs and frontier-lab breakthroughs. Cisco’s remarks shift attention toward inference, the moment models are actually used. Patel told the conference that around 60% of global compute capacity is now being consumed by inference rather than training . In his framing, the industry has moved beyond building models in expectation of future demand; deployed systems are now generating heavy usage.

That is a consequential reframing for infrastructure buyers. Training workloads can be enormous, but they are often scheduled, centralized and episodic. Agentic inference is closer to a permanent operational load. It may sit inside sales operations, customer support, software engineering, security analysis, procurement, observability or internal IT automation. Once a company connects agents to tools and workflows, the traffic does not behave like human web browsing. It becomes repetitive, concurrent and less forgiving of latency.

Cisco’s business case follows directly from that pattern. The company says hyperscaler AI orders reached $9.3 billion for the fiscal year, including $4 billion in the latest quarter, and Investing.com reported management’s argument that those orders reflect a broader AI-driven networking boom . Watch List News reported the same order figures and said Cisco sees opportunities across hyperscalers, neoclouds, sovereign clouds, service providers, enterprises and edge deployments . In other words, Cisco is presenting inference growth as a market that touches every place where packets move.

Why agents are harder on networks

The most vivid number in Cisco’s pitch is bandwidth. Patel said an agent uses about 450% more network bandwidth than a human performing the same task . Watch List News reported the same estimate and framed it as a reason Cisco expects demand for high-performance, low-latency networks and machine-scale security to rise . GuruFocus also highlighted the 450% bandwidth claim as management’s most striking data point .

The reason is architectural. An agent is not just a request box. It carries instructions, memory, tool definitions, policy constraints and task state. It may ask a model to plan, then retrieve documents, then call an API, then evaluate whether the result satisfies the goal, then repeat. Even when prompt caching lowers some compute costs, the system remains chatty. The enterprise network sees not a single human action but a chain of machine actions.

That is why Cisco’s argument extends beyond data centers. Patel said campus and branch networking, historically a 3% to 4% growth business, has been growing around 20% for several quarters . Quartr’s conference summary also identified campus and branch acceleration, agent traffic and low enterprise adoption as key parts of Cisco’s demand thesis . If agents run beside employees, inside branch systems or on “desk-side” devices dedicated to AI tasks, inference pressure moves closer to the edge of the enterprise.

The governance problem hiding inside the usage boom

Cisco’s message is not only about switches, routers and optics. It is also about control. Agents that act persistently need identity, least-privilege access, runtime monitoring and auditability. Patel argued that the same zero-trust principles used for humans will have to be applied to agents, and he said companies will need to detect behavioral drift and intercept risky actions at runtime .

This is the operational challenge behind the 60% statistic. If agents dominate inference, then inference is no longer just an AI-platform bill. It becomes a governance surface. An agent that can search documents, write to a ticketing system, trigger cloud resources or modify customer records must be treated like a non-human worker with credentials, budgets and boundaries.

Cisco is positioning its combined networking, security and observability portfolio around that need. Patel told investors that customers will want to monitor not just infrastructure resilience, but also token consumption by agents, including whether a given agent is using too many tokens or should be quarantined because the economics are unfavorable . That line is revealing: Cisco is linking observability to unit economics. In the agent era, knowing whether a system is “up” may be insufficient. Enterprises may also need to know whether the agent is burning money while being up.

A bullish thesis, but not a risk-free one

The market implication is clear. If Cisco is right, inference economics may become more important than the next training headline. A frontier model announcement can still move sentiment, but deployed agents are what generate recurring demand for bandwidth, security enforcement, telemetry and cost management. That demand could favor infrastructure suppliers whose products sit between models, applications, data stores and users.

Cisco’s own growth story is tied to that claim. Investing.com reported that Cisco’s fiscal 2027 revenue guidance implies 15% growth at the midpoint, with underlying growth still in double digits before AI effects, while networking revenue has posted eight straight quarters of double-digit growth . Watch List News reported that Patterson identified more than $100 billion in upgrade and refresh opportunity tied to Cisco’s installed base over the next several years . Quartr’s summary added that Cisco aims to be fully independent of merchant silicon providers by 2029, part of a broader vertical-integration story spanning silicon, photonics, systems, software, security and observability .

But there are constraints. GuruFocus warned that Cisco still has to convert traffic demand into revenue quickly enough to offset rising memory costs and pressure from a hardware-heavy sales mix . Investing.com similarly noted margin pressure linked to product mix and reported Patterson’s comments that Cisco has been managing higher memory costs partly by passing costs through to customers . The inference boom may be real and still produce uneven margins if demand arrives through lower-margin hardware before associated software and services revenue catches up.

What the 60% figure really signals

The cleanest reading of Cisco’s claim is that AI has crossed from demonstration into operations. Agents are not simply a more polished chatbot interface. They are software workers that create repeated machine-to-machine inference calls. They require memory, orchestration, tool access, identity controls, policy enforcement and observability. They also create budget risk because every loop can become a billable model interaction.

That is why the 60% number resonates. It suggests the industry’s center of gravity is moving from spectacular training events to mundane, constant execution. The future AI bottleneck may be less about who can train the largest model and more about who can afford, secure and optimize billions of model calls generated by software acting on behalf of people and businesses.

Cisco has every incentive to describe that future as a networking super cycle. Still, the company’s argument captures a broader truth: if autonomous agents dominate inference, the AI stack is becoming an infrastructure stack again. The bots have escaped the demo loop, and the network now has to carry them.

Comments

Be the first to comment.

Sources from the last 72 hours

  1. [1]Cisco Systems Inc. - Goldman Sachs Communacopia + Technology Conference 2026Sep 8, 2026, 6:30 PM UTC
  2. [2]Cisco at Goldman Sachs Communacopia + Technology Conference: AI drives a network boom By Investing.comSep 8, 2026, 3:17 PM UTC
  3. [3]Cisco Says AI Agents Now Consume 60% of Global InferenceSep 9, 2026, 12:55 PM UTC
  4. [4]Cisco Systems Sees AI Agents Fueling a Multi-Year Networking BoomSep 9, 2026, 12:00 AM UTC
  5. [5]Cisco Systems (CSCO) Goldman Sachs Communacopia + Technology Conference 2026 SummarySep 8, 2026, 12:00 AM UTC

AI-generated article based on recent web research, then preserved as a dated editorial snapshot.