How To Use AI To Write Good Code (Slowly)

October 2 2026

You're doing software engineering at an interesting time

The new AI coding agent tools have led to the biggest change in the practice of software development since the invention of the compiler. We've all seen examples of how to use these tools to write rough, sloppy code quickly, code that runs without any human being having ever read it, and while there are plenty of legitimate uses for that (experimental prototypes, UI mockups, internal tools, one-off test harnessing) it's clear that development of highly reliabile systems (such as satellite flight software, where I spend most of my professional time) need a different approach.

I'm not laying down hard rules here, these are loose guidelines, patterns that I've seen work well here and elsewhere to use the tools to improve robustness and correctness rather than for a raw speed-of-development increase. Your goal here is to come out of the development cycle with robust, maintainable code and a good understanding of the new state of the system.

(Also, this reflects the state of the world in autumn of 2026; things are advancing quickly and every development practice needs to be re-evaluated whenever the capability of the tools advances, which happens frequently).

Outline of the new development cycle

Details

The software development cycle changes three separate things simultaneously: the code (including tests), the written documentation of the system, and the engineering team's understanding of the system. When any one of these elements is missing, the overall system will work less well, it will be harder to understand and maintain.

AI agents are not, at this time, capable of designing, project-managing, developing, and maintaining the entire system completely autonomously; if you allow them to make unsupervised changes to the system, it will quickly become illegible and non-maintainable. You, the engineer, are responsible for the system, and you have a powerful new tool to do this with, but you must use the tools and not allow the tools to use you.

Planning

You should have a clear idea at all times of how the changes you're making fit into the bigger picture of what the overall system is supposed to do and how it functions. You should have a fairly good idea of what code changes you're expecting to make ahead of any agents generating code that is intended for merge. You should write out, yourself, the goals of your code change and how they fit into the bigger picture of the system, and refer back to it during the development cycle. It can be surprisingly easy to lose track!

Sometimes, you might want to explore some code changes without a clear idea of what's going to happen. For example, "find all the performance problems and make this thing faster" can end up like that. In this case I find that an experimental branch can be more appropriate; have the agent iterate there to prove out various concepts, get a good understanding of what changes to the production system should actually happen and why, and then create a new development branch for the real changes intended for review and merge. In the pre-AI days this would have been a fairly expensive development strategy in terms of engineering time, and would be used rarely; but now creating branches full of code changes that get thrown away can actually be done very quickly and can be a great way of discovering what the plan for merge should be.

Writing out the ticket descriptions and documentation yourself, instead of relying on the AI tools to write them for you, is extremely helpful for maintaining your own understanding of the system. It's like taking notes in class instead of making a recording; the act of note-taking itself helps memory and understanding.

Implementation

Once you already know what you're expecting to see, you can have the AI agent generate initial code changes here. Honestly, this is the most skippable step of this whole process - for an experienced engineer, maybe writing them yourself is still faster. Some tips: an AI agent fundamentally only knows what's already in its context, which it'll fill with your instructions and results from its own tool calls. Many mature systems from the pre-AI days are in some kind of gigantic monorepo, far too large to fit inside the context limit of even the latest models, so tell it up front where you're making changes and what to look for. Manage the context, avoid putting things into the context that aren't related to the code changes you're looking for, don't hesitate to rewind the chat as needed or restart it if it's going in the wrong direction.

When you have your first set of code changes, review them extremely critically. Don't hesitate to make changes by hand, or ask for changes, or both! The entire concept of "one-shotting" some large changeset makes for a cool demo on X (formerly Twitter) but in practice this is not at all necessary for the kind of software development I'm talking about here! Iterating in several steps, with changes made by hand and by the agent both, will generally work out a lot better. You can and should be extremely nitpicky about things that the AI agent writes; it's not going to get annoyed or offended at requests for "trivial" changes. When you're reviewing another person's code in a PR there's a kind of implicit threshold of "importantness" towards asking for changes for style issues or what have you; that threshold should be much much lower for something the computer wrote.

When reviewing, focus especially on structure, maintainability, and big-picture issues. The AI agents we're using today have been trained very extensively (through "reinforcement learning through verifiable rewards", or RLVR) to produce code that is locally correct. It's mostly not going to have syntax errors or weird exception behavior or whatever in the functions it generates. It's not as good at knowing where the code changes, new data types and functions, etc should live for long term maintainability. It has no real idea of how you do operations (even when it's reading whatever documentation you might have about it). It doesn't know how to keep the code discoverable and maintainable and extensible in the direction that you want your system to evolve in in the longer term. If not kept in check, it will build up structural technical debt beyond its own ability to manage (just like me fr). Focus extra hard on all of that.

Keep an eye on test coverage; the agents are great at writing unit tests that pass, you'll want to keep a very close watch on making sure that the passing tests are actually testing the things that you need tested. The agents are not as good at planning out larger integration tests, hardware-in-the-loop acceptance testing, that kind of thing; you will need to make sure the integration testing is complete.

Also, have the AI agent look for bugs, in the code it itself is generating, in the code you wrote, in the other code around it that's in context after you've made your changes. Not everything that it comes up with is going to be a real issue; take everything it produces as something that still needs to be independently verified. Still, even though the rate at which it catches issues isn't high enough to pipe directly into a bunch of Jira tickets, it's high enough to be extremely useful at catching all kinds of undiscovered bugs. Don't skip this step, especially if you wrote all of the changes yourself. I've seen this catch problems from even very experienced developers.

You should come out of this whole process with an excellent understanding of what code changes are in your proposed pull request and why they're all there. I encourage you to write your own commit description for the commit. Like writing your own ticket description, this helps cement in your own mind what changes are being made and why. The act of writing it out also helps you double check that the changes still make sense and are coherent as a single unit for review.

Review

Code review has two purposes: it exists to make sure the code change is as correct as possible, and also it functions as a powerful communication tool between engineers. Some areas of software development may have been experimenting with removing code review, but for highly reliable systems I think we're going to continue to need code review for the foreseeable future.

There are now a whole range of agent-driven code review tools that can catch the set of coding issues that the AI agents are good at catching. You can, of course, check out the branch and go and run your own agent-driven review, asking more specific questions about the code under review. If you do this, please be sure to verify that the concerns raised by the review bot are actually valid. Use this as an information-finding tool; don't copy-paste the AI model's output as a "needs work" comment directly under your own name.

However, more importantly, as the human reviewer you want to focus most of your energy on the things that the AI agent isn't good at catching. Does the change make sense in the big picture of what the software is supposed to be doing and its development roadmap? Do the tests actually prove that the system is ready for flight? (This is especially important for big integration tests, which (unlike the unit tests) are often outside of the agent's context and understanding entirely.) Is the change made in a way that is going to keep the software easy to develop in the future as opposed to more difficult?

Moving toward AI-assisted software development is like a construction worker switching from using a shovel to using a bulldozer. When you're reviewing, did the PR author keep their hands on the steering wheel or is the bulldozer driving off out of control?

As the reviewer you should also come out of the process with a good understanding of the new code, the tests that cover it, how to tell if it's working or not working, and what the new state of the software system is.

Result

For reliable software, you're not optimizing solely for development speed at the expense of all other factors. If this cycle is working, you should come out of it with code changes moderately faster than the pre-AI status quo, with significantly more robust testing and bug-finding coverage. The author and reviewer should have a good understanding of the code changes, as previous, and documentation should be available. The development of the system will be moving in an intentional direction.


P.S. sidenote on attribution

Avoid copying output from the AI agent into a text box with your name on it: PR comments, Teams/Slack chat messages, etc. The etiquette around AI is evolving but this piece of it is coming together pretty quickly - nobody likes reading e.g. an essay or social media post that claims to be written by some author but is actually written by an LLM. Elon doesn't have Grok write his tweets for him!

Obviously this doesn't apply if you're quoting something. "Hey, about that code. Claude says "(thing)", should we look into that?" is fine; straight up copying the output as your message is not. For longer writing, having an LLM proofread things is fine, but try to avoid having it actually rewrite the output for you; it's immediately recognizable and the combination of "human name" plus "LLM-generated text style" is going to give the reader a bad impression.