Klea Merkuri

Klea Merkuri

Sep 29, 2026 · 22 min read

How to Understand AI-Generated Code in 5 Minutes

I’ve spent an absurd amount of time this year building things to make my AI coding agents faster. That means custom skills and MCP servers wired into half my tools.

I also built a whole system for reusing context across sessions, so I’m not re-explaining the same codebase every Monday morning.

Most of it worked. Some of it didn’t, and I ripped it back out a week later.

Now, none of that is what I actually want to tell you about in this post because the best thing I built for myself this year wasn’t designed to make my agent faster at all.

It was designed to make sure I still knew what was happening after it finished 🙈

That sentence sounds backwards coming from someone who likes automation as much as I do. But I built it because I noticed something uncomfortable, and once I noticed it, I couldn’t unsee it.

In this post, I’ll walk you through what it does and how it came about, then provide the skill along with the one user rule I recommend you add to make it all work well.

Explore: What Is MCP, And Why Can’t You Just Use an API?

The Gap Between What an AI Agent Changed and Why It Changed That Way

I’d hand a coding agent a real piece of work to restructure how a form validates before submit, move a piece of state up two levels in the tree, swap out how a config value loads between environments. (The easy examples.)

The agent would finish. Tests were green, and for the most part, the agent attached a summary.

I could tell you exactly what changed, every time. The summary is good at that part.

What an AI Agent’s Completion Summary Doesn’t Tell You

What I sometimes couldn’t do, if you asked me five minutes later, was explain why the agent picked that approach over two or three others that would have worked just as well.

Or how the new piece actually fit into the rest of the data flow. Or what tradeoff it quietly made on my behalf while I was doing something else in another tab 😬

This isn’t a debugging problem. When something breaks, I ask questions. I push back; I ask the agent to walk me through its reasoning, because I have to.

Ordinary implementation work doesn’t force anything like that. The code runs, tests pass, and it’s easy to glance at the summary, nod, and move to the next task.

How many of those sessions had I walked away from without really knowing why the code looked the way it did?

Why This Isn’t About Being a Worse Engineer

I did this dozens of times before I noticed the pattern: I could describe what had happened to my codebase without being able to explain why it happened that way.

That doesn’t mean I stopped being a good engineer.

I think it means the part of the loop that used to teach me something without me trying had quietly gotten delegated along with everything else.

It’s the investigating, false starts, slow crawl through an unfamiliar API that, ironically, all had some value; it had never been totally pointless.

However, I admit that I like the speed too much to want that part back. I like swapping a repetitive refactor for something that demands harder thinking and creativity.

What I don’t like is keeping the speed and losing the reasoning that used to come free with it.

Related: Here’s Why AI Made Developer Burnout Worse, Not Better

The Implementation Brief: A Small Skill for Claude Code, Cursor, and Codex

So I had to do something and ended up building something small. Embarrassingly small, actually, for how much thinking went into figuring out what it needed to not do.

I call it the Implementation Brief. It’s a skill file a coding agent reads and acts on after it finishes a meaningful piece of work. This kind of thing:

  • architecture
  • an API
  • how data flows between two parts of the app
  • a config default
  • a test suite
  • the actual behavior of something a user will see

I built it to run inside agentic coding environments like Claude Code, Cursor, or Codex, rather than locking it to one specific product. Mainly because I work across a few of them.

Related: It’s Remarkable That We Rely On Models We Don’t Own

What I Needed an AI Coding Skill to Actually Do

Before I built anything, I carefully considered what I actually needed, because I didn’t want to write a random skill file and see what happened.

My mental goal wasn’t to replace an AI agent writing code, but to have a way for me to keep learning while using an AI agent to write that code and while I was orchestrating those tasks.

I also had to stay conscious of my running token quota, so whatever I built needed to work with the summary output most agents already hand back, not duplicate it.

The bigger requirement was intent. I didn’t just want an answer to why and how or a list of what had changed in a repo.

I wanted something that helped me stay on top of what was actually done, plus the specific reasoning behind the action itself.

The end objective was knowing the what, how, and why, as well as the intent behind it.

What’s more, it had to be portable because if it only worked in Cursor, or only in Codex, it wasn’t going to survive how I actually work day-to-day, from my professional projects to my personal.

Why I Didn’t Just Use an Existing AI Explanation Skill

Once I had my requirements list, I went looking for something that already did this, hoping I wouldn’t have to build anything at all.

What I found got close in places. Claude Code’s explanatory mode does part of this, as does OpenAI’s PR draft summary skill, along with a handful of handoff skills and code-explanation skills other people have built who’ve run into some version of the same problem.

None of them, however, matched my full requirement.

What the research pointed me toward was something closer to a post-implementation learning contract for coding agents.

One that’s not trying to explain everything that happened because that’s noise that burns tokens for no real benefit. Instead, it should focus on the decisions you should know.

The existing tools proved the pattern was worth building for. But modifying any one of them into what I needed would have been a bigger, unnecessary headache than just writing a small, dedicated skill.

It really helped to establish that the job I needed done was narrow and specific. Kind of like going up to the agent and saying:

“Hey, agent. After a substantive piece of implementation work, produce a compact, evidence-based account of the change, one that lets me understand and communicate the implementation without a separate explanation conversation to get there.”

How the Implementation Brief Decides What’s Worth Explaining

The biggest shift from my original framing that came directly out of the research was that I stopped thinking about the brief as a report and started thinking about it as a filter.

A handful of rules ended up doing most of the work:

Rule 1: Adaptive, Not a Fixed What-Why-How Template

My first instinct was to make what, why, and how three mandatory sections on every single task. That sounds succinct and thorough, but it isn’t 😔

Fix a typo, and now the agent is explaining why it chose that particular string literal. That’s not learning; it’s repetitive AI nonsense, the exact noise I was trying to avoid in the first place.

So the brief had to be adaptive instead of a fixed template.

It never gives a programming tutorial. Instead, it explains the implementation for the specific thing that was actually done, nothing more general than that.

Rule 2: No Exposing the Model’s Chain of Thought

I was also deliberate about the brief not exposing or requesting the model’s private chain of thought.

Instead, it reports engineering rationale grounded in the codebase’s own conventions, the requirements it received, or the actions it actually took.

It doesn’t manufacture a tidy retrospective story about its own thinking after the fact.

Rule 3: Explain Exceptions, Not Conventions

A second rule underneath all that is to explain exceptions, not conventions. Because if the code follows the pattern the rest of the repo already uses, that’s not brief-worthy.

If it breaks from that pattern for a reason, that’s exactly the kind of thing worth a sentence or two.

It’s why the brief preferentially surfaces things like state ownership, component boundaries, API interface shape, error boundaries, or performance implications, and only when one of those actually shows up in the work.

It doesn’t translate the diff into plain English; that’s not the job.

Rule 4: Decisions, Not Activity

All of it comes back to the same central rule: explain decisions, not activity.

A model handing back a summary of activity isn’t the problem since most already do that reasonably well.

But activity doesn’t explain why the agent made a decision, and that’s the actual learning value I was after.

Rule 5: One Filter Question Decides What Makes the Brief

Underneath all four of those rules sits one filter question that the whole skill runs on. It doesn’t ask “what files changed,” but “which 1-3 decisions from this work have the highest learning value.”

I leave everything else out on purpose.

In practice, that means a brief tells me:

  • why the agent chose this approach over the obvious alternatives
  • how the new piece connects to the rest of the system
  • what the tradeoffs were and why
  • what I should actually walk away understanding

It’s short. A few sentences, maybe a short couple of paragraphs depending on the amount of work done, but it’s certainly not an essay.

Note 👇
This doesn’t replace the agent’s usual summary, that’s often useful on its own. Instead, it comes in as a compact learning layer, one that skips restating whatever the summary already covered instead of repeating it just to fill space.

When the Implementation Brief Runs (and When It Doesn’t)

It also knows when to stay quiet. A typo fix, rename, formatting pass, or purely conversational exchange will not trigger a brief.

If there’s no real engineering decision buried in the work, there’s nothing worth briefing me on. The same logic runs in reverse: if I just spent forty minutes chasing a race condition with the agent and found something worth remembering, that’s brief-worthy even when no file ends up changing.

Tip: The test I use is simple. Would future-me, six months from now, benefit from knowing why I built it this way? If yes, brief it; if it’s mechanical, skip it.

Here’s the skill, if you want to look at it or steal it:

---
name: implementation-brief
description: Produce a compact, evidence-based engineering debrief at the end of substantive coding work so a competent developer can understand, explain, and learn from what the agent did without asking follow-up questions. Use before the final response after meaningful code, architecture, state, API, data-flow, configuration, test, or behavior changes, and after substantial debugging or investigation when there is a concrete engineering lesson even if no code changed. Supplement the normal completion summary with learning-oriented rationale without repeating information it already covers. Skip or compress trivial edits, mechanical changes, and conversation-only work.
---

# Implementation Brief

Turn completed engineering work into developer understanding with minimal token overhead. Explain the implementation that actually exists, emphasizing decisions that shaped it rather than narrating every action.

Write for a competent developer. Teach the local implementation, not programming fundamentals.

## Rules

1. **Explain decisions, not activity.** Prefer why a boundary, state owner, API shape, listener, parameter, abstraction, error path, cache rule, dependency, or test strategy has its current shape over a chronological edit list.
2. **Anchor rationale in evidence.** Use task requirements, resulting code, repository conventions, tests, observed behavior, or explicit constraints. Never invent motives or reconstruct hidden chain-of-thought.
3. **Be selective.** Surface only choices that materially affect behavior, correctness, maintainability, extensibility, performance, or future work.
4. **Supplement the normal closeout without repeating it.** Preserve the agent's useful completion summary. Add the brief as a compact learning layer focused only on rationale, mechanics, or takeaways not already conveyed. Do not restate the same facts merely to fill the brief.
5. **Stay compact.** Default to roughly 80-160 words. Use less for simple work and exceed about 220 words only for genuinely independent architectural decisions.
6. **Be concrete.** Name the relevant component, hook, route, state owner, event, query key, API, schema, or test. Avoid unsupported phrases such as "for maintainability."

## Closeout workflow

Before the final response:

1. Inspect the completed work and verification evidence.
2. Pick at most 1-3 decisions with the highest learning value.
3. For each, determine: **what changed, why it has this shape, how the important flow works, and what constraint/pattern made the choice meaningful.**
4. Drop facts that merely restate the diff or obvious syntax.
5. Produce the smallest useful brief.

Never expose private reasoning traces. Report concise engineering rationale supportable from the work.

## What deserves explanation

Prioritize details involving:

- state ownership or synchronization;
- component/module/service boundaries;
- effects, listeners, subscriptions, cleanup, concurrency, or lifecycle;
- APIs, parameters, schemas, types, defaults, or caller contracts;
- reuse of an existing mechanism instead of a parallel abstraction;
- validation, retries, fallbacks, errors, or meaningful edge cases;
- caching, invalidation, query identity, persistence, or data flow;
- dependency/platform choices with plausible alternatives;
- compatibility, performance, security, accessibility, or operational constraints;
- tests that encode an important behavioral contract;
- investigations that rule out plausible causes or reveal a reusable diagnostic technique.

Usually omit imports, renames, formatting, boilerplate, generated code, obvious syntax, and routine commands.

## Code-change closeout

Use **Implementation brief** when a heading helps. Adapt these elements; do not force all of them:

- **Changed** — one sentence describing the behavioral or structural result, not a file inventory.
- **Why this shape** — 1-3 compact points covering the important decisions and concrete rationale.
- **How it works** — a short interaction flow when useful, e.g. `filter -> URL params -> query key -> refetch`.
- **Verified** — concise test/check/reproduction evidence.
- **Worth knowing** — at most one codebase-specific takeaway that helps with the next related change.

For a trivial/mechanical change, use at most one informative sentence or skip a standalone brief.

## Investigation-only closeout

Use **Investigation brief** when substantial debugging produced a useful conclusion but no code changed. Distill rather than repeat the debugging transcript:

- state the conclusion or best-supported cause;
- explain the key mechanism/evidence;
- give one useful implication, ruled-out path, or next debugging boundary when relevant;
- distinguish confirmed facts from remaining uncertainty.

## Quality test

A useful brief should help the developer answer at least one question such as:

- Why is this state owned here?
- Why does this effect/listener need this lifecycle?
- Why does this API or parameter have this shape?
- What is the important data/control flow?
- What existing repository pattern was intentionally followed?
- What would become incorrect or harder with the plausible alternative?

If it cannot, it is probably only a diff summary; improve it or shorten it.

## Examples

### Code change

**Implementation brief**

**Changed:** Added URL-synchronized filtering to the transactions table.

**Why this shape:** The URL remains the source of truth instead of duplicating filter state locally, so refresh/back navigation preserve the view without synchronization code. The implementation also reuses the existing debounced search hook, preserving its cancellation behavior.

**How it works:** filter control -> search params -> route-derived query state -> query key change -> filtered refetch.

**Worth knowing:** Future shareable filters should extend the route schema rather than add component-only state.

### Investigation only

**Investigation brief**

The failure occurs before response parsing: the client builds the endpoint with an undefined account ID. Network handling and JSON decoding are downstream symptoms. It reproduces when navigation happens before account hydration completes, making initialization the next boundary to address. No code changed; the remaining design choice is whether routing should wait for hydration or the client should reject missing IDs explicitly.

Place it at repo level first to give it a try with an active project prior to moving onto a global level. I highly encourage also adding the following user rule:

After substantive coding or debugging work, use the implementation-brief skill to shape the final response. Skip or compress trivial work.

Why an AI Coding Agent Should Brief You, Not Explain Everything

I rewatched What’s Wrong with Secretary Kim a few months ago, a K-drama about an absurdly competent secretary and the executive who can’t function without her. There’s a running bit throughout the show where she walks into his office carrying a stack of folders and orients him on what he needs to know before a meeting.

She doesn’t hand him a full case file and expect him to read it cover to cover. She tells him what matters, in under a minute, and he walks in ready.

That’s the distinction I was actually reaching for with the Implementation Brief, and I didn’t have language for it until I made that connection.

I didn’t want an explanation.

There isn’t time to read a full explanation after every piece of delegated work (let’s be real, when is there ever), and most of the time I don’t need one.

I wanted a briefing.

A full explanation assumes I’m starting from zero. A briefing assumes I’m capable and just need to be oriented before I move on.

It respects that I already know how to think and makes sure I’m not walking away from a session missing the one piece of context that would matter if this code broke at 5 PM on a Friday 🥲

This is also why the brief has to stay short. The second it turns into a wall of reasoning, I stop reading it, and it becomes exactly the overhead I built this thing to avoid.

A briefing that takes longer to read than the change took to review has already failed at being a briefing.

Once I had that framing, I noticed it applies beyond this one skill since a lot of what makes an assistant (human or otherwise) actually useful is knowing what to leave out.

How My Confidence Changed After Weeks of AI-Assisted Coding

Realistically, I didn’t expect the briefs to change much. I added them mostly out of curiosity because it took five minutes to write the skill file after laying down the structure and requirements.

In other words, it was easy to rip out if it got annoying.

But it didn’t get annoying, and after a couple of weeks I noticed I’d stopped asking my agent as many follow-up questions after a task finished. It wasn’t that I’d stopped caring about what happened. The brief was just already answering the question I would have asked (in many cases).

I also noticed I felt more confident walking into the next session. If a teammate asked me why the agent moved a piece of state the way it did, I could provide an answer, instead of going back to reread the diff and reconstruct the reasoning myself.

A few times, reading the brief was the moment I actually disagreed with the agent’s reasoning, not during the work itself. I’d read the short version of why it chose an approach and think, actually, no, that’s not the tradeoff I’d have made here.

Yes, that may be a small thing, one that only happens if I’m engaging with the reasoning, but it’s much better than nothing at all.

Note 💁‍♀️
I ran an experiment on myself, for myself, and it worked well enough that I kept doing it. Whether it makes me faster, or whether it holds up over a year instead of a few months, I haven’t measured. I can only claim that it’s made quite the difference in how confident I feel when working with AI across various problems in my day to day.

I started leaving AI-assisted sessions with a clearer sense of what I’d actually just done. That alone felt like wildly good progress.

Knowledge Debt: The Research That Named a Problem I Was Already Feeling

The research that shaped the skill itself was about tools and requirements. I looked at what already existed, what was missing, and what the brief needed to do differently.

It also revealed the underlying problem was bigger than my own workflow. Unsurprisingly, plenty of people have found themselves in the same situation. (Also explains the existing variety of explanatory skills and tools for agentic workflows.)

The Paper That Coined the Term Knowledge Debt

A group of researchers recently published a paper called Agents That Teach that names this pattern as “Knowledge Debt.” The basic idea is that when AI takes over parts of the work where developers traditionally built understanding through investigation, debugging, and implementation, that learning doesn’t automatically show up somewhere else.

It lines up with research I’d already written about in How AI Makes It Really Easy to Stop Thinking. Taken together, the finding is that when we offload the parts of the work where understanding used to develop, we need to be intentional about putting some of that learning back into the workflow.

Tip 💬
Another way to think about it is that AI does not make developers worse at understanding things so much as it relieves them of the struggle that actually helps them learn. For example, because you learn how to fly a plane, does that mean that you can actually go in the cockpit and fly one? No. And if you do go in the cockpit to fly one (once you have a solid foundation), does it mean that you’re going to have a perfect landing and takeoff on the first try? Most likely not. You’re going to make those slight blunders that you will remember the next time you fly.

With the Implementation Brief, I was trying to make sure that after an agent did meaningful engineering work, I still understood the decisions that matter. I was not trying to solve AI comprehension in general.

The Knowledge Debt framing helped me see that as more than a personal preference. It was a small intervention in the learning loop that agentic development had made very easy to skip.

At the end of the day, the goal isn’t to put the friction back, but to make sure removing the friction doesn’t also remove the learning.

What Self-Improving AI Agents Taught Me About Learning as an Engineer

A lot of what I’ve been building lately, on the AI side, aims to make agents better over time with feedback loops and memory that persists across sessions.

Most recently, that has evolved into reflection steps where an agent looks back at what worked and adjusts. Skills that improve the more they’re used.

In fact, the industry runs on a simple loop: action, feedback, learning, better behavior next time. I kept building that loop for the AI.

Was I building the equivalent loop for myself?

No, and that’s the actual shift the Implementation Brief represents, more than the mechanics of the skill file itself.

My old loop, without me really noticing, had quietly become: prompt, implementation, validation, next task. Nothing in there fed back into what I understood.

The brief adds the missing step. AI-assisted action, briefing, understanding, retained knowledge, better judgment on the next task.

I don’t think engineers need to compete with the agents doing the implementation work. That race isn’t winnable and isn’t worth entering.

The industry is spending a lot of effort on systems that improve themselves through memory and reflection. I think it’s worth spending a little of that effort on workflows that let the humans using those systems keep improving too.

Call it a self-improving engineer, if that helps the idea stick. I don’t love the term any more than I love most AI-adjacent buzzwords.

But the judgment that catches a bad architectural decision, or notices when an agent’s confident answer is wrong, has to come from somewhere. It doesn’t build itself.

Related: I Easily Find Bugs In AI Code Every Single Time

It’s a Wrap

While I don’t have a big system to hand you, I do have a small rule and a small skill that operationalizes it.

Not every AI interaction needs to teach me something. However, the ones involving a real engineering decision should leave me understanding more than I did going in.

That covers architecture changes, new APIs, and state that moves between layers. It covers config defaults, meaningful test changes, and the kind of debugging session where the agent and I figured out something neither of us knew going in.

It doesn’t cover a typo fix, a rename, or a Tuesday afternoon spent reformatting. Nor does it cover something the agent successfully debriefed me on in its response summary (varies by model).

I still use AI for almost everything. In fact, I have to because it’s part of my job description as of late last week. That said, I’m not slower because of the brief, and I haven’t given up any of the speed that made me want to build all this tooling in the first place.

What changed is smaller than speed, and I believe it’s more durable. I’m done trying to optimize the human out of the loop because I’d rather optimize the loop itself—agent and engineer together—so the part of me that used to learn by struggling still gets to learn something, even now that the struggling is optional.

Not everything in an AI-assisted workflow needs to make us faster. Some of it should make us sharper.

What do you think? Let me know, and I’ll see ya next time.

Bye 🙌

😏 Don’t miss these tips!

We don’t spam! Read more in our privacy policy

Related Posts

Leave a Comment

Your email address will not be published. Required fields are marked *