The Bottleneck Is No Longer the Model but the Missing State of the Workflow

--- title: The Bottleneck Is No Longer the Model but the Missing State of the Workflow slug: the-bottleneck-is-no-longer-the-model-but-the-missing-workflow-state meta_title: Why Long-Horizon Agents Need Workflow State meta_description: Zeroset raised $5.2 million for Nebula, a state layer that lets agents query how a company actually routes approvals and handoffs. category: ai tags: Zeroset,Nebula,agent memory,workflow state,Gradient Ventures,long-horizon agents,enterprise AI,process mining,LongMemEval,context ---
A $5.2 million pre-seed is not a large number. The claim attached to Zeroset's is worth more attention than the check: the thing blocking enterprise agents is not model capability but the absence of a system of state.
Akshat Kannan left Stanford at 18. William Zhang left the University of Texas at 21. They founded Zeroset in January 2026 around a narrow observation. Frontier models already handle email drafts and flight searches. They stall when a task requires knowing how a specific company routes approvals, exceptions and handoffs.
On October 6, Gradient Ventures and 2048 Ventures co-led a $5.2 million pre-seed, with Leblon Capital participating. The company has five people, and access remains a closed research preview.
What Nebula stores
Nebula connects to the systems a workflow already uses: Microsoft 365, SharePoint, Outlook, Teams, GitHub and workplace messaging. It does not treat each document as an isolated vector. It builds a hierarchical vector graph with three layers: semantic facts, episodic event clusters and procedural habits.

The distinction from a memory store is what the agent gets back. Rather than reconstructing history, it queries the current state of a workflow. The company describes the return as typed, queryable state with lineage: what is true, when it became true, and why the system believes it.
The benchmarks are early but specific. On LongMemEval, Nebula returned 20 percent higher accuracy than Mem0, Supermemory and naive RAG, at a median under 50 milliseconds and with fewer tokens per query. On the BPI Challenge process-mining datasets, it beat all recorded baselines on next-activity prediction by more than 10 percent, which is a harder test than retrieval because it asks whether the system can anticipate the next legitimate step.
The token count matters as much as the accuracy. Fewer tokens per query means a smaller context window can carry the same information, which lowers cost on every agent turn. For a workflow that runs for a week, that difference compounds the same way the errors do.
The founders' starting observation is the tightest statement of the problem. Frontier models handle email drafts and flight searches. They stall when a task requires knowing how a specific company routes approvals, exceptions and handoffs. Nothing about the model is wrong in that scenario. The information simply is not in the prompt and cannot be retrieved from a public corpus.
The integration surface is the real moat
Connecting to Microsoft 365, SharePoint, Outlook, Teams, GitHub and workplace messaging is unglamorous work, and it is also the hardest part to copy. A competitor can replicate a three-layer graph design in a quarter. Reproducing the permission-aware connectors that keep a state layer consistent with what each user is allowed to see takes far longer, and enterprise buyers evaluate it first.
Permissions are the part that most memory products handle badly. A workflow state layer has to respect the same access boundaries as the underlying systems, so an agent acting for one user cannot surface a fact that only becomes visible at another level of the organization. Getting this wrong does not produce a wrong answer so much as a compliance incident.
It also explains why the company is running a closed research preview rather than a self-serve product. State layers are only as good as the workflows they have seen, and every new customer integration surfaces edge cases in routing logic that no benchmark covers.
What buyers should ask before adopting
Three questions separate a useful state layer from a demo. Does the system record when each fact became true, and can it invalidate a fact that has been superseded? Can it explain why it believes something, tracing back to the source event? Does it enforce the same permissions as the systems it reads from?
The lineage requirement is the one that tends to be answered vaguely. A state layer that returns a fact without provenance is functionally a cache, and a stale cache inside an autonomous agent is worse than no memory at all, because the agent will act on it with confidence.
Persistence is the second-order question. Continuous state across days or weeks is the entire value proposition, and storing that much derived organizational knowledge raises a data governance conversation about retention, deletion and regional residency.
Why stale state compounds
The failure mode the company names is the reason the category exists. A stale fact or contradictory instruction does not cause a single error. It becomes the starting state for the next turn, producing compounding rework, extra tool calls and human intervention.
The compounding is what makes it expensive. A single wrong action is recoverable. A wrong assumption that persists through a twenty-step process corrupts every step after it, and by the time anyone notices, the work has to be redone from an earlier checkpoint. Agents that run unattended for hours make this worse, because nobody is watching the drift.
Conversational memory systems retrieve similar chunks. If the retrieved fact is three weeks old and the workflow moved on, the agent reasons from a world that no longer exists, and nothing in the response flags the discrepancy. Retrieval similarity and factual currency are different properties, and a vector store optimizes for the first.
Nebula's three-layer graph is an attempt to give the agent a way to tell them apart. Semantic facts describe what is true. Episodic event clusters describe what happened. Procedural habits describe how the organization usually does things. Lineage is the piece that makes it auditable: the fact, when it became true, and why the system believes it.
Where this fits
Mem0 and Zep occupy the adjacent memory-infrastructure category. Zeroset's bet is that long-horizon enterprise agents need process state rather than conversational memory, and the distinction matters for anything spanning days or weeks across system boundaries: claims handling, release approvals, multi-step procurement.
Those processes share a structure. They involve handoffs between people and systems, exceptions that follow unwritten rules, and approvals that depend on who is available. None of that lives in a document. It lives in practice, which is why it has been hard to represent.
Process mining is the closest prior art, and it is telling that Zeroset benchmarks against process-mining datasets. That field spent two decades reconstructing how work actually flows from event logs. Doing the same thing in real time, so an agent can query it mid-task, is a reasonable evolution.
Gradient's published thesis makes the same argument. Consumer agents succeeded because they trained on the public web. Enterprise agents still lack a representation of how work actually moves inside a private organization.
What changes if it works
Two things. Agent harnesses stop shipping ever-larger context windows to compensate for missing structure, because the problem was never window size. And enterprises can keep the operating model of their own workflows inside their own boundary, rather than hoping a general model eventually infers it.
The second consequence is the more strategic one. A workflow representation that lives inside a company's boundary is an asset the company owns. One inferred by a general model is a capability it rents. For regulated industries, that difference decides whether an agent can be deployed at all.
There is a competitive risk worth naming. If a major platform vendor decides process state belongs in its suite, a five-person startup faces a distribution problem regardless of how good the state layer is. Gradient's involvement helps, and distribution is not something a $5.2 million round solves.
The numbers today are small: five people, two public benchmarks, a closed preview. The claim is larger, and it is testable. If enterprise automation has moved past the capability stage, the next round of competition is about who knows what the work actually is.
Related articles
Satellite Photos Are Now Robot Training Data
The bottleneck in physical AI training stopped being compute or model capability. It became the quality of the synthetic world.
Google's Gemini 3.5 Live Translate Removes the Pause
Translation that runs continuously, in the speaker's own voice, on a phone already in your pocket, moves the feature from something you open to something simply on.
China Wrote the First Mandatory Safety Standard for AI Agents
Safety moves from a feature you advertise to a gate you pass. The risk inventory sits at 13 categories and 97 items.
AI-Generated Content Now Has to Declare Itself
This step doesn't solve every problem, but it turns “AI-generated” from an option you could hide into a question you have to answer.