One Telos

Why One Telos works on agent reliability

Ali Shah ·

Until now, One Telos offered a 14-day, fixed-price engagement to find a company’s best AI opportunity. It was a reasonable test, but it sat in the most crowded part of the market. Every agency now sells AI strategy.

From today, One Telos works on one problem: making software that acts on its own do what it was meant to, fail safely, and leave evidence of what it did. AI agents are the current case. The problem gets bigger as the models get better, not smaller.

Where agents stall

Confluent’s 2026 survey of 4,625 IT leaders found about a third running agents in production. Among those, more than three quarters report stalled projects, and roughly two thirds of all respondents name LLM reliability and non-determinism as a barrier.1

The teams that do ship keep their agents on a short leash. A study of 306 practitioners found that most production agents run ten steps or fewer before a human steps in, that most teams rely mainly on human evaluation, and that reliability is the leading development challenge.2 The market is not asking for full autonomy. It is asking for bounded agents it can depend on.

Mostly engineering failures

When an agent fails in production, the model is rarely the cause. In February 2026 an n8n upgrade produced tool schemas that both OpenAI and Anthropic rejected, and enterprise workflows stopped until teams rolled back.3

Schema drift, lost state, duplicated side effects and silent loops are distributed-systems problems. They have known answers: idempotency keys, outboxes, sagas and durable execution. Those answers are becoming the foundation of agent platforms. Replit, OpenAI’s Codex web agent and Cursor run long-running agent work on Temporal,4 and Temporal’s OpenAI Agents SDK integration reached general availability in March 2026.5 It is event sourcing and sagas under a new name.

I spent eight years on these problems with Akka, Kafka, Kinesis and DynamoDB Streams, before the systems in question called a language model. The failure modes have not changed much. The consequences have, because the component making decisions is now non-deterministic.

Observability is unsettled

OpenTelemetry moved its GenAI conventions to a separate repository in June 2026. As of late August they were still in Development, with no tagged release.6 Vendors are consolidating: ClickHouse bought Langfuse in January 2026, and Cisco announced its intent to buy Galileo.7 The practical advice is to instrument against the standard and treat the vendor as replaceable.

A clock

Under Regulation (EU) 2026/1744, the EU AI Act’s Annex III high-risk obligations now apply from 2 December 2027. The Article 50 transparency duties kept their 2 August 2026 date, and the Commission treats agents as covered by the Act’s existing definitions.89

I am not a lawyer, and what the Act requires of a given firm is a question for its legal team. The engineering work underneath is easier to state: being able to trace what a system did, with which inputs, and why. That is observability work, and fifteen months is not long to add it to a system that was not built for it.

What One Telos does now

One Telos helps engineering teams take AI agents from pilot to production, and prove they do what they’re meant to, fail safely, and leave an audit trail. The first offer is the Telos Review, a 14-day, fixed-price reliability audit. Alongside it I am building the Failure Lab: a refund-processing agent, broken on purpose and then hardened, with the numbers published.

The thesis, not the tools

Each wave of computing, from mainframes to the web to cloud, has needed people who make it dependable. The more autonomy software gets, the more that matters. Today the material is agents, evals and durable execution, and it will be swapped out as the field moves. The question underneath it will not: does this system do what it was meant to?

Telos is Greek for purpose, or intended end. The International AI Safety Report defines reliability as the degree to which an AI system works as its developer or user intended.10 That is the work.

If you have an agent that worked in the demo and misbehaves in production, I would like to hear about it. Book a 20-minute call or email a.shah@onetelos.com.

References

  1. Help Net Security, on Confluent’s 2026 Data Streaming Report, 18 June 2026.
  2. “The Hierarchy of Agentic Capabilities”, summarising Pan et al. (2025). arxiv.org/pdf/2601.09032
  3. Data Science Collective, “Why AI agents fail in production”, April 2026.
  4. Reactify Solutions, “Durable AI agents in 2026”, 12 June 2026.
  5. Jacar, “Durable agent execution with Temporal”, July 2026.
  6. TrueFoundry, “OpenTelemetry GenAI conventions”, 22 August 2026.
  7. G. Anhaia, “OpenTelemetry GenAI semantic conventions”, DEV Community, 2026.
  8. Praxikon, on the Digital Omnibus and Regulation (EU) 2026/1744.
  9. AlpacaX, “EU AI Act high-risk deadline moved to 2027. What didn’t?”
  10. International AI Safety Report 2026. arxiv.org/pdf/2602.21012