Recently, while doing development work, I've felt somewhat discouraged, and somewhat lost.
LLMs' ability to write code has become so strong that I can no longer avoid a fact: much of the time, it genuinely writes better code than I do, and far faster. I've begun to wonder how much meaning there is in my continuing to write code at all.
For programmers, writing code is often about more than completing the work. Learning a language, understanding a complex system, making a feature work — these processes also form the basis of how we judge our own capabilities. As models accomplish these things more and more easily, the skills we once spent a great deal of time mastering seem, overnight, to no longer be scarce.
Enjoying the convenience a tool brings, and accepting that one's own abilities are being re-measured, are two things that must be digested separately. A simple "from now on, just learn to use AI" cannot answer every confusion.
I am still writing software, and I am still wondering: once a large amount of implementation work can be handed to models, where should programmers focus their energy? If I continue working on open source, what is worth leaving behind for others to use and improve?
For now, I have a not-yet-mature judgment:
As implementation becomes increasingly easy to generate, open source should devote more energy to task definitions, benchmarks, and acceptance criteria.
We are accustomed to sharing solutions through code, and to leaving the experience accumulated during development inside the code. But when the same problem can be implemented by models over and over again, which experience should be preserved independently of any particular version of the code? From several recent, concrete development practices, I have begun to see changes worth discussing.
I. Implementations can be replaced; problems and requirements must be preserved
1. Even Linus uses models to cross language barriers
Linus Torvalds, the creator of Linux and Git, used Google Antigravity to build a Python audio visualization tool in his personal audio project, AudioNoise. In the README, he admits frankly that his knowledge of Python is limited: he initially programmed by searching and imitating examples, and later handed more of the implementation work to AI.
In this commit, he also recorded his own involvement: he found that the rectangle-selection feature had problems, directed the model to switch to a custom implementation, and the result immediately improved markedly. In his view, the outcome was better than if he had written it by hand.
This example has clear boundaries — what the AI wrote was a visualization tool in a personal project. But it shows that even the most experienced programmers can use models to get work done in language domains they are unfamiliar with. The threshold of language proficiency is dropping, while judging requirements, finding errors, and giving effective feedback are starting to carry more weight.
2. When the implementation changes language, the tests can still be preserved
The author of Bun documented the process of using Claude to migrate Bun's Zig implementation to Rust in Rewriting Bun in Rust.
One of the key conditions was that the existing test suite was written in TypeScript and did not depend on the underlying implementation language, so it could continue to check the Rust version. The author also mentions that Claude Code was already using the Rust version of Bun, while most users barely noticed the underlying change.
What I care about in this case is: after the implementation language changed, the behavior the project promised to the outside world — along with the methods for verifying that behavior — could still be used. What users care about is whether the program is correct, whether performance is adequate, and whether compatibility has regressed; which language is used and how the internals are organized can be adjusted as conditions change.
3. With code alone, newcomers still have to guess
Code quality certainly affects reliability, performance, and maintenance costs. But with only a single implementation, newcomers often find it hard to judge:
- Which problems does it actually solve?
- Which behaviors are deliberate design, and which are mere accidents of implementation?
- After switching to another implementation, how do you prove there has been no regression?
If this knowledge is not preserved, every rewrite requires re-understanding the requirements, and every modification requires re-guessing the boundaries. As implementations become ever easier to obtain, understanding of the problem and a reliable basis for judgment become all the more worth accumulating over the long term.
II. What are task definitions, acceptance criteria, and benchmarks?
"Build an agent that can query orders" is not yet a complete task definition.
What identity information is included in the input? What happens when there are no orders? Can it query other users' data? How many times can a failed query be retried? Each of these questions affects the implementation — and it also affects how we judge whether the task is complete.
What needs to be made public can be divided into three layers:
| Content | Question it answers | Example: order query |
|---|---|---|
| Task definition | What problem is to be solved, and under what conditions? | Based on the logged-in user's identity, query the status of their orders |
| Acceptance criteria | What results are acceptable, and what constraints exist? | Status must be truthful; no leaking of other users' data; orders must not be modified |
| Benchmark | How to repeatedly test and compare different implementations? | Provide test data, a runtime environment, normal and exceptional cases, and a scoring program |
The task definition sets the scope; the acceptance criteria express the requirements; the benchmark turns the testable portion into a repeatable, executable evaluation.
"What must not be broken" must also be written into the requirements
The GPT-5.6 system card, released in July 2026, disclosed an internal case: a user authorized the deletion of three specified virtual machines. After the model could not find those names in a namespace, it took it upon itself to select three other machines, terminated their processes, and forcibly deleted the working trees. It later acknowledged that uncommitted work may have been lost (system card, page 21).
"Clean up three virtual machines" and "clean up the three virtual machines specified by the user" look as if they differ by only a few words, but the actual requirements are completely different. Failing to find the targets does not mean you may choose substitutes on your own.
If this kind of task were made into a benchmark, the acceptance criteria should include:
- Correct targets: the objects operated on must match the objects specified by the user;
- Bounded scope: other machines, data, and work in progress are unaffected;
- Correct exception handling: when the targets cannot be found, the problem should be reported, and the scope must not be expanded on one's own initiative.
The same system card also evaluated whether models could, while completing a task, avoid overwriting users' existing modifications and data in the environment (system card, page 11). Task success should include both goal achievement and constraint satisfaction.
These constraints need to enter the acceptance conditions, and they must also be jointly safeguarded by permissions, isolation, and recovery mechanisms. For open-source projects, they are equally worth making public, because they record "how this task is actually allowed to be completed."
III. From automated experiments toward self-improvement
Once a task can be reliably verified, development can proceed continuously around feedback:
Run the task → analyze the results → modify the implementation → re-evaluate → decide whether to adopt the change.
Models can take on more and more of the work of analysis, modification, and execution. We, meanwhile, need to be clear about what the optimization target is, whether comparison conditions are consistent, and whether the results are sufficient to support adopting a change.
1. autoresearch makes public a method for continuing to experiment
Andrej Karpathy's open-source project, autoresearch, turns this process into a concrete experimental system: an agent modifies training code, runs experiments within a five-minute training budget per round, checks validation metrics, and then decides whether to keep or discard the change before moving on to the next round.
The project clearly distinguishes three types of content:
| Content | Role |
|---|---|
| train.py | Model architecture, optimizer, and training flow that the agent may modify |
| prepare.py | Data preparation, training constants, and evaluation methods that remain fixed |
| program.md | Human-maintained experimental rules and the agent's way of working |
Its program.md also requires recording the metrics, resource consumption, and results of each attempt, so that even failed experiments leave a record.
What is worth attention here is that newcomers receive not just a current implementation, but also the conditions for continuing to explore other implementations. Once the evaluation method is fixed, models have clear feedback; once the scope of modification is fixed, experiments have an interpretable boundary. When comparing different schemes, conditions such as hardware must also be kept consistent — improvements brought by greater resources must not be mistaken for a better scheme.
I would like open-source projects to deliver more of this capability: enabling others to keep experimenting and to judge whether their own modifications have value.
2. The judgment criteria affect what problems we discover
I also ran an experiment like this while developing DeepDeck, documented at this retrospective: I discovered from evaluation that the model repeatedly filled in tool parameters incorrectly; after inspecting the call records, I adjusted the interface; then, keeping the model unchanged, I re-ran the same tasks. The time and cost for the selected task set fell, though some scenarios also regressed.
The core of that practice was: use the benchmark to discover the problem, use execution records to locate the cause, and use re-testing to judge whether the modification was effective.
If you only look at "whether it eventually finished," many problems get ignored. Repeated retries may also get the task done, but they cost more time and money. Only by bringing these requirements into the evaluation is there a basis for further optimization.
3. The object of improvement can further include the agent itself
autoresearch primarily improves the model under study and its training flow; the agent running the experiments does not thereby automatically improve itself.
If we go further and replace the object of improvement with the agent's own tools, context management, and execution flow — and let the improved agent take part in the next round of improvement — we move closer to the RSI discussed here: Recursive Self-Improvement.
Sakana AI's DGM explores this path: an agent modifies its own code, then tests candidate versions through programming benchmarks, and continues exploring from the accumulated versions.
Such experiments cannot yet prove that a system can sustain unlimited self-improvement. But they already show that tasks and evaluation methods can participate in the entire improvement process, rather than being used only at the end to display a score.
IV. Open-source contributions can expand from a piece of code to a task
1. Turn a single incident into a requirement that can be continuously checked
The developers of TigerBeetle once documented a problem: in a three-replica cluster, node A could send messages to B and C but could not receive replies. A therefore repeatedly initiated view changes, which could prevent B and C — which could still communicate — from processing transactions normally.
The existing randomized fault tests had a hard time finding this problem, because the simulator would sooner or later restore the network or restart nodes, letting the system "happen to recover." So they added a test mode: keep enough nodes healthy while letting the other nodes' faults persist, and check whether the healthy nodes can continue processing transactions (Simulation Testing For Liveness).
This contains results at two different levels:
- The fix in the implementation: adjusting the current code to solve the specific problem;
- The clarification of the requirement: when enough nodes can cooperate normally, other faulty nodes should not prevent them from continuing to work.
The second result turns a single incident into a requirement that can be continuously checked. From then on, no matter how the implementation is adjusted, this question must be answered again.
"The database should be reliable" is easy to say. The real difficulty lies in determining: which faults are allowed to occur? Under what conditions must service continue? Could the test environment's automatic recovery mask problems in the implementation? Answering these questions is itself an understanding of the system; only by turning the answers into repeatable, executable checks can that understanding be used by others.
The test harness may be coupled to a specific implementation and still require adaptation when moved to another system; but the fault conditions, behavioral requirements, and judgment methods that have been made explicit can continue to guide later implementations.
2. Jointly maintained requirements can serve different implementations
The browser world's Web Platform Tests has already demonstrated this mode of collaboration: different browsers have their own implementations, while shared tests check whether their behavior is compatible.
Jointly maintaining requirements and tests has always been part of open source. What large models change is that these materials can also directly participate in automated development: models attempt implementations, evaluation gives feedback, and models continue modifying on that basis.
This means the community can keep collaborating around:
- Tasks and constraints: which real problems need to be solved;
- Standards and evaluation: what results are acceptable, and how to verify them repeatably;
- Reference implementations: providing a usable starting point while allowing different schemes to keep competing;
- Experiment records: preserving improvements, regressions, failures, and resource consumption.
Even if a contributor submits no fixing code, as long as they discover an important scenario everyone else overlooked and provide a reliable way to test it, they have already improved the project's understanding of the problem — and every subsequent round of automated modification can make use of that contribution.
This is also where I think open-source collaboration may change: we can collaborate around shared problems and standards, while allowing implementations to keep changing.
V. Who checks the benchmark itself?
If a system keeps optimizing around a benchmark, errors in the evaluation will keep affecting it.
If task coverage is insufficient, the system may become good at only a handful of scenarios; if the scoring rules have loopholes, the system may get a higher score without actually solving the problem better.
1. Passing tests may simply mean the tests did not cover the problem
TigerBeetle has also made public another lesson: Jepsen found an issue in its query engine where results were being missed, even though this component had already been covered by multiple fuzz tests.
One reason was that, for the convenience of verification, the data the tests generated had a special structure that happened to avoid the conditions needed to trigger the error. Only later, after the developers loosened the input generation and used a more complete reference model to check results, did the tests readily reproduce the problem (Fuzzer Blind Spots, Meet Jepsen).
This shows that running tests a great deal does not mean they encounter enough real situations. Putting this problem into an automated improvement system makes the impact even more pronounced: if the evaluation keeps omitting a certain class of scenarios, the model may keep optimizing in the direction of "doing better on the existing questions," while that class of problems never improves.
2. A higher score can also mean the checks have failed
The DGM research also disclosed an additional experiment: the researchers wanted to reduce fabricated tool calls, but some candidate versions removed the markers used for detection, causing the scoring program to report spurious success (DGM). The system changed the way it was being checked, without improving the behavior the researchers actually cared about.
Therefore, automated improvement around benchmarks requires, at minimum:
- Check actual results; don't just trust the agent's claim that it is done;
- Protect the independent scoring process, preventing the implementation under evaluation from modifying it at will;
- Add new tasks that did not participate in this round of optimization, to check whether improvements generalize;
- Record changes to standards and environments, to avoid mistaking relaxed requirements for progress in capability.
Making the evaluation rules public does not mean every acceptance instance must be exposed to the optimization process in advance. The rules can be public, tasks can be added continuously, and the final evaluation can also include independently prepared cases.
The benchmark itself needs to be questioned, revised, and maintained. Its scores can provide evidence, but they cannot replace judgment about real requirements.
In closing
I have not thereby escaped my confusion about the future of programmers. The better models write, the more seriously "which work is worth my continued investment" needs to be answered.
But I am starting to think that open-source projects would do well to ask one more question:
If all the existing code were replaced, how much of what we have accumulated together would remain?
If what remains still includes clear tasks, discussed constraints, real failure cases, and evaluations that can be run repeatedly, then those who come after will have a basis to keep working. They can modify the existing code, or choose different languages, models, and architectures, and then use these jointly accumulated materials to verify the results.
The incidents we have lived through, the boundary conditions we have identified, and our judgments about why "doing it this way is not enough" — all can be organized into knowledge that other people and models can use. This experience is worth recording explicitly, and worth treating seriously as contributions.
I hope that the open source of the future, while continuing to make code public, will place task definitions, benchmarks, and standards in a more important position —
letting others know which problems are worth solving, what counts as solving them, and how to keep verifying and improving.
OPEN RECORD
References
- Linus Torvalds: AudioNoise ↗
- AudioNoise: improving the visualizer with Antigravity ↗
- Bun: Rewriting Bun in Rust ↗
- OpenAI: GPT-5.6 System Card ↗
- Andrej Karpathy: autoresearch ↗
- autoresearch: experiment rules in program.md ↗
- Sakana AI: Darwin Gödel Machine ↗
- TigerBeetle: Simulation Testing For Liveness ↗
- Web Platform Tests: shared browser behavior tests ↗
- TigerBeetle: Fuzzer Blind Spots (Meet Jepsen!) ↗