← Back to blog
AiAbout 6 min read

Where Enterprise AI Actually Pays Off: A Steelmaker and an Airline

Published Oct 4, 2026
Where Enterprise AI Actually Pays Off: A Steelmaker and an Airline

Two case studies landed this month, and together they say more about enterprise AI than a year of vendor keynotes. One is a steelmaker that deployed hundreds of specialized agents. The other is an airline that cut contract review time in half. Neither involved a general-purpose assistant answering questions in a sidebar.

Three hundred agents at a steel plant

Tata Steel, according to Google Cloud, deployed more than 300 specialized AI agents in nine months. They handle asset maintenance, information access, and faster customer responses. The word doing the work in that sentence is "specialized." This was not one agent given a broad mandate. It was three hundred narrow ones, each tied to a specific operational job.

That is a different architecture than the one most vendors sell. A general agent that can do anything is easy to demo and hard to trust. A fleet of agents that each do one thing is tedious to build and much easier to govern, because the scope is defined by the task. When one of them misbehaves, you know which process it belongs to and what it should have been allowed to touch.

The maintenance angle is worth noting on its own. Asset maintenance is where industrial AI has the clearest payoff, because the data is already there, the failure modes are expensive, and a ten percent improvement in uptime is measurable in currency. It is unglamorous work, and it is exactly the kind of work that survives a budget review.

A contract review that went from days to hours

Cebu Pacific reported that its ChatGPT Enterprise rollout cut legal contract-review time by roughly half. Project-intake turnaround fell from five days to about one, and engineering reference searches dropped from ten to fifteen minutes down to under ten seconds. In a survey of users, 86 percent said AI supported more than a third of their daily work.

The contract number is the one to dwell on. Legal review is a task with real stakes, a defined input, and a clear output. It is also a task where the bottleneck is reading, and reading is something language models are genuinely good at. The airline did not automate the lawyer. It removed the part of the job where a person scans a document for the clauses they have seen a hundred times. That is a workflow change, not a tool purchase.

The pattern underneath both

Neither case study is about a model that can reason better than a human. Both are about models inserted at a specific step in a defined process, where the before-and-after can be measured. Tata Steel knows which maintenance decisions got faster. Cebu Pacific knows how many days the intake process used to take.

The number that should worry vendors is in the survey data. Singapore's enterprise AI maturity index put adoption of agentic AI at 51 percent, up from 22 percent the year before, while only 10 percent of enterprises had redesigned end-to-end workflows around it. A Capgemini study found similar gaps, with most organizations still exploring or running pilots.

So the market is full of companies that have connected AI to a workflow without changing the workflow. They bolted a model onto the old process and called it transformation. The two case studies here did the other thing, and that is why they produced numbers worth reporting.

Specialization is what makes the agents governable

There is a governance reason to prefer a fleet of narrow agents over one broad one, and it is worth spelling out. An agent with a fixed scope is easy to audit. You know what it should be able to reach, you can log what it actually reached, and the gap between the two is a short list. An agent with a broad mandate has a scope that changes with every task, which makes the same audit a research project.

Tata Steel's 300 agents are specialized in exactly this way. That choice pays off twice. It is easier engineering, and it makes the deployment defensible to the people who have to sign off on it. Each agent can be reasoned about on its own, and a failure in one does not require reexamining the other 299.

The same principle explains why so many enterprise agent programs stall. A general-purpose agent shows promise in a demo, gets piloted against a broad slice of work, and then runs into the fact that nobody can describe what correct behavior looks like across that whole slice. Narrowing the scope feels like a step backward and is usually the step that makes the project survivable.

The hidden requirement: a baseline to compare against

Both case studies produce numbers, and that is not an accident. Measuring a ten percent uptime gain or a drop from five days to one requires a before. Organizations that skip the baseline end up with anecdotes instead of results, which is why so many AI programs report enthusiasm and struggle to report impact.

The discipline starts before deployment. Pick the process, record how long it takes and what it costs today, then change one step and measure again. It sounds bureaucratic and it is the difference between a program that can defend its budget and one that cannot. Cebu Pacific can say its contract review time was halved because someone wrote down what it used to be.

There is a second measurement that matters just as much and gets taken less often: the error rate. An agent that halves the time and doubles the mistakes has not improved anything, it has moved the cost from the front of the process to the back. The case studies that hold up are the ones where a human still checks the output at the point where a mistake would be expensive, and the measured gain comes from removing the reading, not the reviewing.

What the gap means for the next year

The organizations seeing results are the ones that picked a process, measured it, and rebuilt it around what the model is actually good at. The ones stuck in pilots are the ones that bought access and waited for adoption to happen on its own. Access was never the hard part. Redesign is.

The steelmaker and the airline also share a trait that is easy to overlook. Both kept humans at the decision points that carry liability. The maintenance agent flags, a person decides. The contract agent drafts and highlights, a lawyer approves. The AI removed the scanning, not the judgment.

That division of labor is unremarkable in hindsight. It is also the thing that separates an AI program that produces a number from one that produces a slide. Pick the process, find the step where a person is reading or moving information, automate that step, and measure what changed. The recipe is simple. Executing it is harder than buying a seat, which is why so many organizations buy the seat and call it a strategy.

Related articles