---
title: "Claude writes the playbook. Ballet runs it."
description: "Claude is great at writing a process. It is not a place to run one. We keep hearing the same question from Ops leaders who are not engineers: can't we just build this in Claude? The short answer is no. Not if you need the same refund at 2am, a person who can run the process but not rewrite it, and a password that never lives in a chat."
canonical_url: "https://docs.ballet.dev/articles/claude-writes-the-playbook-ballet-runs-it-odiAydjDXG"
md_url: "https://docs.ballet.dev/articles/claude-writes-the-playbook-ballet-runs-it-odiAydjDXG.md"
---
# Claude writes the playbook. Ballet runs it.

**Claude is great at writing a process. It is not a place to run one. We keep hearing the same question from Ops leaders who are not engineers: can't we just build this in Claude? The short answer is no. Not if you need the same refund at 2am, a person who can run the process but not rewrite it, and a password that never lives in a chat.**

We've been in those rooms.

Someone has watched Claude draft a refund reply. Sketch a CRM update. Write the rule for which queue a ticket should go to. The demo works. The conclusion arrives fast: we can just build this in Claude.

I understand why. You describe the work in English and something useful comes back. For a first try, that feels like the whole product.

It isn't.

The people asking this are not naive. They have a queue. They have a CRM. They have been told, for eighteen months, that AI can do the work. What they have not been shown is the difference between an AI that can _describe_ a process and a system that can _run_ one, the same way, every time.

A chat is a demo. Operations is the eighth identical run. It is the person who should be able to issue the refund and not be able to change the refund window. It is the Salesforce password that must never live in a transcript.

Claude is not God. Claude is an extraordinary writer. Ballet is what you run after the writing is done.

## This is not just our sales story

I went looking at what the rest of the industry is publishing. The pattern is the same.

Gartner said in June 2025 that **more than 40% of agentic AI projects will be canceled by the end of 2027**. The reasons they named were not "the model is too weak." They were cost, unclear business value, and weak risk controls. Their analyst, Anushree Verma, said most of these projects are hype-driven pilots, and that **many jobs sold as "agentic" do not need an agent at all**. Use an assistant to look something up. Use a saved process for routine work. Use an agent only when a real decision has to be made.

That is the same split I am asking Ops teams to make.

S&P Global asked more than a thousand companies already investing in AI. In 2024, 17% said they had abandoned most of their AI projects before those projects went live. In 2025, that number was **42%**. On average they threw away almost half of their proofs of concept.

![image.png](https://docs.ballet.dev/api/attachments.redirect?id=77d14ab3-4658-4537-9fcb-a219d0e165f1 " =960x540")

MIT's NANDA group looked at custom, enterprise-grade GenAI tools. **60% of organizations evaluated one. 20% reached a pilot. 5% reached production.** Chatbots get tried because they are easy. They stall in real workflows because they do not remember your process and they do not own it.

![image.png](https://docs.ballet.dev/api/attachments.redirect?id=4fa88725-f364-4d0e-b923-cf76cf2b83f0 " =960x540")

Claude's own maker draws this line in public. Anthropic calls a **workflow** a system where the path is already written. They call an **agent** a system where the model picks the path as it goes. Their advice: start simple. Use a saved path when the work is known. An agent costs more, and mistakes compound.

Kong, which sits in front of APIs for a large share of the Fortune 500, says the same thing in engineering language: capture a success once, turn it into something you can run again, stop improvising.

I am not claiming we invented this. I am claiming Ballet is the product that does it for Ops: Claude writes the playbook. Ballet runs it.

## The same answer, every time

Ops teams do not get paid for a clever first try. They get paid for the same decision, on the same ticket, when nobody is watching.

This is not a Ballet talking point. In 2025, researchers published a public test of customer-service agents that have to talk to a user, follow a policy, and update a database. Even a strong model (GPT-4o) got retail work right about **61%** of the time on one try. On airline work, about **35%**. When they asked the same retail job **eight times**, the chance of getting it right every time fell **under 25%**.

One good run is a demo. Eight identical runs is a workflow.

![image.png](https://docs.ballet.dev/api/attachments.redirect?id=996371e8-e3df-4a85-a4d0-7c09ac109f17 " =960x540")

That test is not Ballet versus anyone. It is the industry's own evidence that "it worked in the chat" is not the same as "it will work on Tuesday's queue."

We tested the same idea on everyday Ops work: refunds, overdue invoices, tax rates, discounts, license seats, ticket priority. 54 real cases. Each case, eight times. We already knew the right answer.

A Ballet playbook gave the same answer **54 of 54** times. Typical time: **less than a tenth of a second**.

When we asked Claude the same 54 jobs, it stayed consistent on **42 of 54**. Another leading model stayed consistent on **35 of 54**. Both were about **40 times slower**.

|           | Same answer all 8 times | Typical time |
| --------------- | ----------------------- | ------------ |
| Ballet playbook | **54 of 54**            | 0.08s        |
| Claude          | 42 of 54                | 3.70s        |
| Other model     | 35 of 54                | 4.15s        |

![image.png](https://docs.ballet.dev/api/attachments.redirect?id=19bb69f9-0f5b-4ba0-a694-be8b4bc9516b " =1280x720")

At 10,000 similar tickets a month, asking Claude again costs about **$113**. Asking the other model again costs about **$67**. The saved playbook does not pay Claude again. You still pay a small Ballet bill to run it. You do not pay Claude on every ticket.

![image.png](https://docs.ballet.dev/api/attachments.redirect?id=a9507324-7bc7-4487-b1d5-f78e3333739c " =1280x720")

Think of a refund. The rule is already known: under 30 days, approve it. 31 to 45 days, only if the item is unopened and they are on a paid plan. After 45 days, send it to a person. Tax treatment changes for EU customers. Escalations go to a different queue.

If that rule lives only in Claude, Claude has to re-read the ticket and re-decide, every time. On the 200th ticket of the month, the day-30 cutoff can get fuzzy. If that rule lives in a playbook, the playbook reads the order date, the plan, and the region, and walks the same steps. No new guess. No new policy.

The full tables are in [Deterministic operations, measured](/articles/deterministic-operations-measured-A0mAp58RxY). I am not saying Ballet is smarter than Claude. I am saying: write the rule once. Stop asking Claude to invent it again on every ticket.

That is the split we built the product around.

**Use Claude to write the playbook. Use Ballet to run it.**

## "We'll just build it in Claude"

The other version of the same question is money. Not "can Claude write this?" It is "why would I pay Ballet when I already have Claude?"

Gartner's forecast is useful here too. They said projects die on cost and unclear value, not on whether the model can draft a reply. A do-it-yourself Claude setup looks cheap until someone has to host it, watch it, handle retries, and fix it when Freshdesk changes.

We looked at a month of real customer usage and compared the first-year cost of one workflow, a Freshdesk auto-reply at 2,238 runs a month, against building the same thing yourselves on Claude.

Ballet: **$3,303**.

Build it yourself on Claude: **$13,873**.

**Seventy-six percent cheaper**, and live in days rather than weeks.

![image.png](https://docs.ballet.dev/api/attachments.redirect?id=6955442d-2f0c-463c-9f57-9f61f88c9ed1 " =960x540")

On that workflow, a Ballet run costs **9 cents**. That 9 cents is the real bill to own and run it. It is not the same number as "we did not pay Claude again" in the test above. Both matter. They answer different questions.

The year-one numbers are in [Ballet vs Claude code production cost benchmarks](/articles/ballet-vs-claude-code-production-cost-benchmarks-PoDd09T6dX).

## Who is allowed to change the rule?

Would you let your sales team update your CRM from Claude?

Who can change deal stages? Reassign accounts? A rep says, "Move this deal forward," but there is no next step. Your rules should catch that, however they ask.

When the chat is the system, those questions have nowhere to live. Anyone in the thread can talk Claude into a different policy. Tuesday's instructions are not Monday's instructions. There is no saved version. There is no "this person can run it and that person cannot edit it."

This is also what the security research is warning about. IBM and Ponemon's 2025 breach report found **13%** of organizations had a breach of an AI model or application. Of those, **97% had no proper AI access controls**. **63%** of the breached organizations had no AI governance policy, or were still writing one. Shadow AI, tools people use without approval, added as much as **$670,000** to the average breach cost. The average breach they studied cost **$4.44 million**.

![image.png](https://docs.ballet.dev/api/attachments.redirect?id=e473242a-fb62-44c8-afd2-d92167cbb04d " =960x540")

The 97% is not "97% of all companies." It is 97% of the companies that already had an AI-related breach. That is worse, not better. The ones who got hurt were the ones who had not put a lock on who the AI can act as.

Practitioners writing about CRM agents say the same thing in plainer language. Do not give Claude an admin password. Do not let a model inherit a user's full Salesforce rights. Approve refunds, stage changes, and anything a customer can see. Log who asked, what ran, and what changed.

Ballet splits those jobs.

Owners and editors write the playbook. Everyone else asks and runs it. Your team talks to an assistant that collects the inputs, runs the process, and explains what happened. It cannot change the steps. It cannot invent a new refund window because the conversation felt urgent.

You choose who can use which playbook. It is not "anyone who found the chat." The rule lives in a saved playbook, not in a prompt that drifted. When work has to pause for a human, it continues on the same run. Same history. Not a new chat that forgot what already happened.

I am not saying we have built every Salesforce permission checkbox, or a full approval product. I am saying the governance question has a home: who can change the process, who can only run it, and what happened afterward.

That is the RevOps question I keep coming back to. When your team can update the CRM through AI, where do your rules live?

## Passwords do not belong in a chat

The security problem with "just use Claude" is quieter than a wrong refund. Someone pastes a password into the chat so the demo works. That password now lives in history. The chat has become a key ring.

IBM's number is the industry version of that story. If the AI can act, and there is no access control, the breach is not a model problem. It is a keys-in-the-chat problem.

```mermaid
flowchart LR
  subgraph bad [Just use Claude]
    paste[Key lives in the chat]
  end
  subgraph good [Ballet]
    vault[Key stays in the workspace]
  end
```

In Ballet, passwords stay in the workspace. The person running a playbook is never asked to paste one. Apps connect through a normal login, not a secret in a paste. The work that talks to Salesforce or Freshdesk happens in a locked room, not inside Claude's reply.

When a step would do something you cannot undo (issue the refund, close the ticket, write the CRM), the run can stop and wait for a person. When they continue, it is the same run, not a second try that might charge the customer twice.

Retries are an Ops problem, not an "ask Claude again" problem. If the system hiccups, it may try again in a way the other system can ignore as a duplicate. If the playbook itself is wrong, it fails once. That is how you keep a retry from becoming a second refund.

## What Ballet actually is

Ballet is the operations layer next to Claude.

You describe the work. Claude helps you write a playbook: the refund window, the tax rate, which queue this goes to. That playbook then runs as a saved process: same steps, same checks, a record of what happened.

People can talk to that process in Studio. They can ask in plain language. Talking to it is not the same as being the process. Claude can be a brilliant way to draft the playbook. It is a poor place to keep the playbook, the permissions, the passwords, and the history.

Two jobs, on purpose. Writing can stay creative. Running stays boring. Boring is the feature.

## What I am not saying

I am not saying Claude is useless. I use it to write playbooks.

I am not saying Ballet never varies. If a step asks Claude again while it is running, or if it has to call Salesforce live, that part can still change. The "same answer 54 of 54 times" number is from a controlled test, not a promise about every customer workflow.

I am not saying we beat Claude at thinking. 54 of 54 versus 42 of 54 means: did we get the same answer again? It does not mean who is smarter.

I am not saying the industry charts above are Ballet results. They are other people's measurements. We are using them as context, not as a scoreboard.

If this is the conversation you are having with your Ops team, that everything can just be built in Claude, I want that conversation. Ask me for a demo.

## Sources

Industry figures in this piece come from public research, not from our lab:

* Gartner, [Over 40% of agentic AI projects will be canceled by end of 2027](https://www.gartner.com/en/newsroom/press-releases/2025-06-25-gartner-predicts-over-40-percent-of-agentic-ai-projects-will-be-canceled-by-end-of-2027) (25 June 2025)
* S&P Global Market Intelligence, Voice of the Enterprise: AI and Machine Learning, Use Cases 2025 (1,006 respondents). Reported in [CFO Dive](https://www.cfodive.com/news/AI-project-fail-data-SPGlobal/742784/)
* MIT NANDA, [The GenAI Divide: State of AI in Business 2025](https://fortune.com/2025/08/18/mit-report-95-percent-generative-ai-pilots-at-companies-failing-cfo/)
* Yao, Shinn, et al., [τ-bench](https://arxiv.org/abs/2406.12045), ICLR 2025
* Anthropic, [Building effective agents](https://www.anthropic.com/engineering/building-effective-agents) (19 December 2024)
* Kong, [Deterministic AI Architecture](https://konghq.com/blog/engineering/deterministic-ai-architecture-enterprise-reliability)
* IBM / Ponemon, [Cost of a Data Breach Report 2025](https://newsroom.ibm.com/2025-07-30-ibm-report-13-of-organizations-reported-breaches-of-ai-models-or-applications,-97-of-which-reported-lacking-proper-ai-access-controls)

Our own numbers remain in [Deterministic operations, measured](/articles/deterministic-operations-measured-A0mAp58RxY) and [Ballet vs Claude code production cost benchmarks](/articles/ballet-vs-claude-code-production-cost-benchmarks-PoDd09T6dX).
