The demo is always a Tuesday. Clean tenant, rehearsed prompt, an agent that reads an invoice, matches it to a PO, and posts it while the room nods. The count on the slide is enormous. Oracle points to more than 600 agents in Fusion. SAP pointed to 2,500 Joule skills in April, then stood on stage in May and reframed the story around 50 assistants and 200 agents. Gartner says 40% of enterprise apps will carry task-specific agents by the end of this year, up from under 5% in 2025.
The cutover is always a weekend. And the cutover weekend does not care how many agents were on the slide.
| The headline | What it actually counts | The number underneath |
|---|---|---|
| 2,500 Joule skills | Reframed on stage two months later as 50 assistants and 200 agents | When the headline moves that far that fast, it is measuring the pitch, not the capability |
| 600+ Fusion agents | Mixes agents and assistants across Fusion and the industry applications | 22 genuinely agentic applications when they were named in March. Two percent of the headline |
| Thousands of vendors claiming agentic AI | Gartner named the pattern: agent washing, relabeling assistants, bots and RPA as agents | Around 130 judged to be building the real thing |
The count is a marketing artifact
Start with the numbers, because they are doing more persuading than informing.
When SAP's own headline moves from 2,500 skills to 200 agents inside 2 months, the count has stopped measuring capability. It is measuring the current shape of the pitch. Oracle's 600-plus mixes agents and assistants across Fusion and its industry applications. The genuinely agentic tier, the pieces that reason and act rather than retrieve and summarize, was 22 applications when Oracle named them in March. That is a real and interesting number. It is also 2% of the headline.
Gartner named the pattern directly: agent washing, the relabeling of assistants, bots, and RPA as agents. Of the thousands of vendors claiming agentic AI, the firm judged only around 130 to be building the real thing. So when a slide says hundreds of agents, live now, the first job is translation. How many act without a human. On your kind of data. In production, not early access.

What "agent" means in each vendor's mouth
The word carries three different meanings depending on who is selling.
Sometimes it means a copilot: it drafts, suggests, and summarizes, and a person accepts or rejects. Useful, low risk, and not autonomous in any sense that changes your controls.
Sometimes it means a workflow with a language model bolted on: it follows a fixed path and calls the model at a step or two. Reliable, because the path is hard-coded, and barely more "intelligent" than the automation you already run.
Sometimes it means the real thing: the agent decides what to do next and acts on systems that move money or records. That is the version worth the excitement and the caution. It is also the smallest category on offer, and the one vendors are quietest about pricing and reliability for.
None of these is bad. Confusing which one you are buying is what gets programs in trouble.
The reliability number the vendors do not lead with
Salesforce, an agent vendor with every reason to flatter the category, published a benchmark on its own technology. On single-step tasks, its agents succeeded around 58% of the time. On multi-step tasks, the kind a real process actually is, success fell to about 35%. The same study found near-zero built-in awareness of what data should stay confidential.
Read that as a buyer. A capability that completes a realistic multi-step task one time in three is a genuine assistant and a dangerous unattended worker. That gap is why, in the field, every consequential agent action still routes back to a human for confirmation. The autonomy is real in the demo and conditional in production. McKinsey's 2025 survey lands in the same place: 39% of companies reported any enterprise-level earnings impact from AI, and most of that was under 5%. The technology is working. The unattended, bet-the-close version is mostly not here yet.
What still breaks at cutover, agent or no agent
Here is the part the demo never shows, because none of it improved because of AI.
Data conversion still runs past its window. On a statewide program this year, one benefits conversion ran so far past its slot it nearly pushed go-live and blocked the first payroll run. One workstream, and the whole date was at risk. No agent on the slide touches that.
Security still stalls testing. I once watched an end-to-end cycle stop for a full week because one security role exposed compensation data far wider than intended. Every test script that touched pay had to pause. The build was fine. The design decision was not.
Integrations are still a program inside the program. Discovery lists 50. The real number lands closer to 90 once you map every handshake to banks, carriers, tax engines, and downstream apps. Agents add integration surface. They do not remove it.
The cutover weekend is a test of your data, your sequencing, and your governance. AI changes the demo. It does not change the weekend.
Three questions that expose roadmap-ware
Bring these to any agentic demo. They take the conversation from theater to fact in about 5 minutes.
One. Show me this agent acting on our data, not yours. Ask them to run it on a copy of your messiest master, not the clean tenant. If the answer is "in a later phase," you are looking at roadmap, not product.
Two. What does it do with no human in the loop, and what is your published success rate on multi-step tasks? A confident vendor has a number. A vague one is selling a copilot wearing an agent's coat.
Three. What consumes billing when this runs at production volume, and where do I watch the meter? If the person in the room cannot answer, the meter is running somewhere you cannot see, and that is the subject of a later article in this series.

How to run a pilot that produces a real answer
An agent doing something impressive once tells you nothing about the pilot. It is a success because it answered a question you wrote down first. Define the transaction, the data it runs on, the success threshold as a number, the human who owns the exceptions, and the credit budget with a hard stop. Then let it run on real volume. A pilot with those five things produces a decision. A pilot without them produces a testimonial, and you cannot deploy a testimonial.
Do this week: before your next agentic demo, write down the single transaction you would actually let an agent run unattended, and the error rate you would need to see to believe it. Bring that sentence to the demo. Make the vendor meet it on your data, not theirs.
Sources
Oracle, "Oracle Introduces Fusion Agentic Applications," Mar 24, 2026 (22 new agentic applications), oracle.com; Oracle Fusion Insider, AI World, Oct 15, 2025 (600+ agents and assistants), blogs.oracle.com.
SAP News, "SAP Business AI Release Highlights Q1 2026," Apr 2026 (2,500+ Joule Skills), news.sap.com; SAP Sapphire, "SAP Unveils the Autonomous Enterprise," May 12, 2026 (50+ Assistants, 200+ agents), news.sap.com.
Gartner, task-specific agents in 40% of enterprise apps by 2026 (Aug 26, 2025), gartner.com; over 40% of agentic AI projects canceled by end-2027, agent washing, ~130 real vendors (Jun 25, 2025), gartner.com.
Salesforce AI Research, CRMArena-Pro, arXiv:2505.18878, May 2025 (~58% single-turn, ~35% multi-turn agent success; near-zero confidentiality awareness), arxiv.org. McKinsey, "The State of AI in 2025," Nov 2025 (39% report any enterprise-level EBIT impact), mckinsey.com.
Field anecdotes from 9Nation program experience. Clients anonymized; the statewide program is a statewide public-sector Workday program.
