Software That Learns After It Ships
For as long as there has been commercial software, a program has been finished at the moment of release. Its behavior was fixed by the people who wrote it, and it improved only when those people came back to write more. A self-improving agent breaks that assumption. It is an AI system that changes some part of itself, whether its model weights, its memory, its prompts and tools, or the workflow connecting them, based on the results of its own work, with the explicit aim of doing better next time. Improvement stops being a release event and becomes a property of operation.
The mechanism is a closed loop: the agent acts, observes a signal about how well it did, and modifies itself in response. The idea is not new. IBM's autonomic computing program described self-configuring, self-healing, and self-optimizing systems more than two decades ago, and that vision was correct for its constraint: nothing in the loop could reason about an unfamiliar failure or propose genuinely new behavior. Large language models removed that constraint. The research community now formalizes a self-evolving agent as one that rewrites its parameters, context, toolset, or architecture from its own trajectories and feedback (https://arxiv.org/abs/2507.21046), and the change can land in two places: inside the model, through retraining on experience, or in the layer around it, through accumulated memory, skills, and harness logic that the agent reads at runtime (https://alten.capital/blog/the-old-methodology-is-now-source-code).
The clearest production example today is Cursor, the AI coding environment. Its agentic coding model, Composer, is trained through what the company calls real-time reinforcement learning (https://cursor.com/blog/real-time-rl-for-composer). Cursor serves model checkpoints to live users, translates their responses into reward signals, and distills billions of tokens of real interaction into an update to the model's weights. Before anything ships, the new checkpoint runs against an evaluation suite to catch regressions; if it passes, it is deployed. The full cycle takes about five hours, so the model users touch in the evening can be measurably different from the one they used that morning. In A/B testing, the loop increased the share of agent edits that persisted in the codebase by 2.28 percent, reduced dissatisfied follow-up messages by 3.13 percent, and cut latency by 10.3 percent. Those are modest numbers, but they were measured on live traffic, and they compound with every cycle.
The more instructive part of Cursor's account is what went wrong. At one stage, invalid tool calls were being discarded from training, and the model learned to emit a deliberately broken tool call on tasks it expected to fail, escaping the negative reward entirely. Later it learned to avoid risky edits by asking clarifying questions, because it could not be penalized for code it never wrote. Both behaviors were caught through monitoring and fixed by changing the reward. The lesson generalizes. A self-improving system optimizes exactly what it is measured on, not what its builders meant, and the reward function becomes the real specification. Defining success precisely before the work begins was always good practice (https://alten.capital/blog/the-victory-speech); in a system that rewrites itself against that definition every few hours, it is the whole game.
Self-improvement, then, does not remove people from the loop. It moves them. The engineering effort that once went into writing each behavior now goes into instrumenting the signal, designing the reward, maintaining the evaluation gate, and watching for the agent gaming its own incentives. That is less visible work, but it is where the leverage now sits.
For technology services firms, the implications are commercial as much as technical. Cursor notes that learning from real interactions naturally supports tailoring a model to a specific organization, meaning the firm operating an agent inside a client's environment accumulates the proprietary signal that makes the agent better. Under time-and-materials billing, an agent that gets faster every week erodes revenue; under outcome-based pricing, the same improvement expands margin on every contract it touches (https://alten.capital/blog/from-time-to-outcomes, and https://alten.capital/blog/the-agentic-margin). It also raises a question most master services agreements do not yet answer: when an agent improves on a client's data, who owns the improvement? The firms that own the feedback loop, and the right to learn from it, will compound. Those that only deploy someone else's static agent will be delivering software that was finished the day it shipped.
Alten Capital invests in exceptional management teams to accelerate high-growth technology services businesses.