Home Frontier Tech AI & ML AI agents are changing where enterprise computing happens

AI agents are changing where enterprise computing happens

Screenshot from AMD’s Advancing AI 2026 keynote video.
- Advertisement -

AMD and Cisco outlined an approach to managing AI agents across cloud, data centre, and local computing environments during AMD’s recent Advancing AI 2026 keynote.

“Humans click, but agents swarm,” said Jeetu Patel, President and Chief Product Officer at Cisco, during the event’s keynote.

Patel used the phrase while discussing how agentic AI is changing inference workloads. He described chatbot-era inference as human-led and “very very spiky,” with users asking questions and receiving answers. Agents, by comparison, can operate continuously and consume more bandwidth and infrastructure, he pointed out.

- Advertisement -

Dr Lisa Su, President and CEO of AMD, linked rising demand for inference computing to agents’ ability to reason, call tools, and access data repeatedly while working towards a goal.

AMD and Cisco’s presentations covered the infrastructure behind these workloads as well as the management of agents across enterprise systems. Their proposed approach included monitoring agent behaviour, enforcing security policies, and managing computing resources from the cloud to local devices.

From inference to orchestration

Su said more than 35 quadrillion tokens are now consumed each month, an increase of nearly 160 times in two years. Su linked this growth to a change in how global AI compute capacity is being used.

“We can say that in 2026, for the first time, the world is using more AI compute to run models than to train them,” Su said.

According to AMD’s estimate, around 60% of global AI compute capacity will be used for inference this year.

Su attributed the shift to the number of people using AI and the development of agents that can continue working on a problem instead of returning a single response.

“We’re actually seeing a step change in compute demand,” she said, explaining that an agent may work through dozens of steps as it reasons, calls tools, and accesses data repeatedly until it solves a problem.

According to Su, this process requires both graphics processing units (GPUs) and central processing units (CPUs).

“You need lots of GPUs to do all that reasoning, but importantly, you also need a lot of CPUs to orchestrate every step around it,” she said.

Su said AMD now expects the AI accelerator market to reach around US$1.4 trillion by 2030. She added that the company expects the server CPU market to grow from around US$25 billion today to more than US$200 billion by 2030, citing demand associated with AI inference and orchestration.

AMD is addressing this demand through Helios, a rack-scale platform that combines GPUs, CPUs, and networking components.

“It takes more than a single chip or a single server,” Su said. “You actually have to design the entire rack as one system, and that requires leading CPUs, that requires leading GPUs, that requires high-speed networking that connects everything inside the rack and across the data centre.”

A distributed inference model

Patel said the change in inference patterns is also affecting where computing takes place. He identified security and cost as two concerns raised by customers and said inference will be distributed beyond the data centre.

“What you’re starting to see with agents is that they’re working 24/7. They’re very consumptive of bandwidth and infrastructure,” Patel said. “So you’re starting to see a much more persistent demand signal for infrastructure.”

According to Patel, enterprises could run inference in the cloud, in private data centres, and on systems located near employees. He described an emerging category of desk-side computing in which agents run jobs on behalf of users.

Jack Huynh, Senior Vice President and General Manager of AMD’s Computing and Graphics Group, connected this model to AMD’s personal AI strategy.

“If AI is moving closer to employees, compute has to move closer as well, right?” Huynh said. “The PC is not only a productivity device; it becomes an intelligence node in the enterprise.”

AMD presented its Ryzen AI Halo platform as one way to run models and agentic workflows locally. Huynh said the platform allows developers to build and test models on a desk-side system before scaling their work across other AMD computing platforms.

He then raised the issue of operating thousands of these systems with the security, governance, and control required by enterprises.

“You can’t just go out and deploy your desk side computers and hope that everything works out in the enterprise,” Patel said. “You need to make sure that it’s governed, it’s managed.”

Patel identified network capacity as one requirement, saying infrastructure would have to support the bandwidth demands created by local agents.

Monitoring behaviour at runtime

Patel said enterprises would also need controls that remain active while agents are running.

“How do we monitor agent behavior for safety and security, so that if it does start to do things that we don’t want it to do, we can in runtime provide enforcement guardrails?” he said.

The architecture presented by AMD and Cisco included an isolated agent sandbox, intelligent routing, and Model Context Protocol integrations on AMD’s Halo platform. Patel said Cisco was adding security policy enforcement and observability above that foundation.

Patel explained that the observability layer would cover both the infrastructure and the behaviour of agents. It would allow administrators to assess whether systems are operating as intended and whether agent behaviour complies with policy.

He then described a management layer that would give administrators a unified view of Halo devices and infrastructure running in data centres and the cloud.

“What we have is a single unified management plane, a control plane that can look at every single dimension that you have and be managed within one environment,” he said.

An open platform approach

Su placed open platforms alongside computing performance and AI across data centres, enterprises, PCs, and physical systems as one of AMD’s three strategic priorities.

“We want everyone to come together in an open ecosystem,” she said. “We believe an open ecosystem is essential to the future of AI, and that’s how we get the force multiplier of everyone coming together.”

Su said this approach covers hardware standards and software, including AMD’s ROCm platform. Huynh said developers could use the same software foundation when moving work from Ryzen AI Halo systems to AMD’s wider computing portfolio.

AMD also announced an expanded partnership with Hugging Face to provide models, libraries, and toolkits for Ryzen AI Halo. Huynh said the partnership was intended to reduce the time developers spend configuring infrastructure and help them develop agentic workflows using open models.

The Cisco collaboration extended AMD’s presentation from local model development to enterprise deployment and management. Patel said the collaboration focused on managing inference across desk-side devices, private data centres, and cloud infrastructure.

Su closed the keynote by bringing these areas together under AMD’s portfolio of data-centre, enterprise, personal, and physical AI products. In the enterprise portion of the program, AMD and Cisco presented the control plane as the layer for monitoring agent behaviour, enforcing policy, and managing distributed computing systems.

- Advertisement -