Skip to main content
← ArticlesLire en français
4 August 2026·Alien6 Research

When the AI agent starts forgetting what it already knows

With Watchdog, Alien6's R&D explores how to measure, contain and improve the work of development agents without sacrificing quality or turning their governance into a surveillance system.

LLMArchitectureProduction

With Watchdog, Alien6's R&D explores a question that is still barely visible: how do you measure, contain and improve the work of development agents without sacrificing quality or turning their governance into a surveillance system?

The turning point did not come from the monthly bill. It appeared during a session that had grown so large that the agent was spending a growing share of its work retrieving what it already knew.

The files it had read, the test results, the tool outputs and the delegated searches had gradually piled up. The model had more information available, but that did not necessarily make it more relevant. The context, meant to help it, was starting to become a burden.

This situation captures one of the paradoxes of agentic AI: the more memory and autonomy we give an agent, the more we need to be able to observe how it uses them.

Watchdog was born from that experience — an R&D and experimentation project carried out at Alien6 around Claude Code.

Watchdog does not claim to provide a definitive answer to agent governance. It is first and foremost a testing ground. Its goal is to understand real usage, identify drifts and test optimisation mechanisms before considering their generalisation.

Context, the debt you never see

A conversational agent does not memorise a mission the way a human would. At every new step, it must retrieve the useful information from a context made of the conversation, the project instructions, the files it has read and the results produced by its tools.

This working memory has a cost. The larger it grows, the heavier interactions become. A session sometimes ends up containing stale results, several reads of the same resource or the parallel searches of different sub-agents.

Intuition would suggest that a larger context necessarily makes the model perform better. That is not always the case. The useful information can end up drowned in noise. Cost rises while the quality of the answer stagnates, or even degrades.

The problem is therefore not just about picking a cheaper model. You have to determine which level of intelligence to mobilise, with how much information, and for which task.

Watchdog starts by making these phenomena visible. Its dashboard breaks down costs by project, session and model. It tracks how context evolves, distinguishes the different forms of cache and measures the share consumed by sub-agents.

An unusual expense then stops being a mere line on a bill. It becomes a technical signal. It can reveal an overly long session, excessive delegation or a workflow that deserves rethinking.

But measuring after the fact is not enough.

Watchdog can step in when a drift appears. It can flag that a context is getting too large, recommend compaction, suggest starting a new session or limit certain redundant delegations.

It can also shrink a very long technical output while keeping the useful errors, prevent the re-reading of an unchanged resource or point out that an elementary command could be executed directly.

Asking a model loaded with tens of thousands of tokens of context to carry out a simple operation is sometimes like mobilising a team of architects to move a chair.

These controls rely as much as possible on deterministic rules. Watchdog does not call a second AI to notice that a file is too large or that an image has already been analysed. Governance must not become yet another layer of consumption.

From observation to experimentation

The most interesting dimension of the project appears when the data starts to produce lessons.

Watchdog can analyse locally the recurring usage patterns of a project: code exploration, implementation, debugging, testing, documentation, operations or log analysis. It is no longer only about knowing how much a session cost, but about understanding what was actually done.

A regularly repeated workflow can reveal the value of a specialised skill. A frequent summarisation task can be handed to a local model, within a strictly bounded frame. The recurring use of an external service can reveal the need for a dedicated integration.

These recommendations are not applied silently. They are documented, previewed and left to the user's decision. When an optimisation is activated, it opens an experiment.

Watchdog then keeps a sample of reference sessions and observes the new sessions matching the same activity. The cost per turn can be compared before and after the change, but it is never interpreted on its own.

This caution matters. A cheaper session may simply have been shorter. It may also have involved fewer files, fewer turns or a smaller initial context. Concluding too quickly that an optimisation is effective would amount to confusing correlation with causation.

To better address this difficulty, a local AutoML mechanism has just been introduced into Watchdog.

Watchdog turns observed usage into measurable recommendations, with a local AutoML experimentation loop.

The term may sound ambitious. Here, it deliberately refers to a narrow, controllable setup. When enough sessions are available, Watchdog compares several statistical models. The first measures the apparent effect of the optimisation. The next ones attempt to correct the analysis by accounting for the number of turns and the size of the context.

Cross-validation then selects the model that produces the smallest errors on observations it did not use to fit itself. In other words, Watchdog does not select the explanation that best describes the past, but the one that seems to hold up best when confronted with new data.

The goal is not to produce a spectacular prediction. It is more modest, and probably more useful: reducing the risk of attributing to an optimisation a saving that actually comes from less complex sessions.

The result always comes with an uncertainty interval. It is presented as an association observed before and after the change, never as automatic proof of causation.

AutoML therefore does not decide in the user's place. It helps choose the analysis method best suited to the available data.

A saving only has value if quality holds

The experimentation does not stop at measuring costs.

Watchdog also examines several quality signals: test results, corrections needed after validation, stability of the changes or delivery signals. An optimisation cannot be recommended if it reduces consumption while noticeably degrading the work produced.

This is what distinguishes Watchdog from a mere FinOps dashboard. Saving tokens is not an end in itself. What matters is the ratio between the resources consumed and the value actually created.

The same caution applies to the data used to establish this diagnosis. Watchdog runs locally. To understand usage, it keeps derived signals such as activity categories, the tools invoked or certain quality indicators. The user's requests and the content of the files examined are not copied into its analysis database.

The project thus seeks to observe the system without watching the developer.

This distinction becomes essential as agents gain autonomy. Observability must not produce an opaque apparatus of individual control. It must allow the user to understand what the agent is doing, why it is doing it and which resources it mobilises.

An applied research approach at Alien6

Watchdog reflects the way we approach R&D at Alien6.

The value of an innovation does not lie solely in the power of a model or the sophistication of its architecture. It appears when that innovation can be turned into a system that is understandable, measurable and usable in real conditions.

The project thus sits at the intersection of software architecture, applied AI, FinOps, developer experience and statistical analysis.

It also allows appealing ideas to be confronted with reality. Should a task systematically be delegated to a sub-agent? Does a local model actually produce a saving once the coordination cost is taken into account? Does reducing the context improve the next session? From how many observations can a reasonable conclusion begin to be drawn?

Watchdog does not yet provide a definitive answer to all these questions. That is precisely its role as an experimental project: making hypotheses explicit, building the necessary instruments and accepting that some intuitions will be contradicted by the data.

This approach is part of the work carried out at Alien6, at the intersection of strategic IT consulting, software architecture, artificial intelligence and applied research.

It also matches the way we tackle these subjects: starting from an observed difficulty, building a system able to measure it, then testing the possible answers without hiding their limits.

AI agents will progressively take on longer sequences of work. Their autonomy will grow. The need to understand their consumption, their decisions and their limits will grow with it.

Tomorrow, the question will probably no longer be just which model to use. It will be how to make sure that model properly mobilises the resources, the tools and the context entrusted to it.

Further reading

LLMs in Production: 5 Architecture Patterns

15 October 2025

Why Your RAG Pipeline Is Underperforming

15 December 2025

Related services

aiarchitectureengineering