Karpathy's New Advice: Ask AI to Build a Tool for Understanding

Andrej Karpathy's new post opens with a prediction: as language models take on more work, we will spend more time understanding their outputs. His suggestions are concrete. Ask for clearer language, a diagram, an interactive webpage, or even an explainer video made for an arbitrary topic. Original post
What interests me is the additional job this gives an agent. We already ask agents to research, write code and prepare proposals. We can also ask them to build something that helps the recipient understand, question and try the result.
That changes the conversation after delivery. Faced with a long report in an unfamiliar field, I might once have said, “Explain it again.” Now I can ask, “Draw this relationship,” “Give me a page where I can change that condition,” or “Animate the process so I can follow it.” Helping me understand can become part of the assignment.
What Karpathy is proposing
The post moves through writing, diagrams, webpages and videos. Each step introduces a richer form; bespoke explainer videos are the format Karpathy is especially optimistic about. Original post
There are two ideas here. One is his experience: constraining the writing and changing its form can make model output easier to understand. The other is his forecast: greater autonomy will move human effort toward oversight and understanding, while increasingly abundant code makes custom explanation tools more plausible to create.
Together they suggest a useful direction: after an agent delivers the work, keep asking it to help you get closer to that work. If the answer feels distant, ask it to build a bridge. The bridge might be a diagram, a page or a video.
The post does not present a complete agent-management system or a comparison of learning outcomes. What follows is my practical extension of its suggestions: start where understanding gets difficult, then ask the agent to change the handoff.

Original concept illustration: four forms with different jobs, without a measured ranking.
First: make the writing name actions and conditions
Karpathy's first trick is to ask for an explanation in ASD-STE100. This controlled technical English was developed to help people understand aviation maintenance documentation. He reports finding its constraints useful, sometimes asking for a looser version that gets about eighty percent of the way there. Original post, official ASD-STE100 background
The useful idea is to constrain language. Give the same object the same name throughout. State the actor, action and condition. Split a complicated sentence. Avoid letting an impressive but vague word replace a specific behavior. We can request these principles in Chinese too; that does not turn Chinese prose into compliant technical English. A model accepting a prompt is also different from a standards audit.
Suppose an agent writes, “A robust idempotency mechanism ensures consistent retry behavior.” It sounds professional. You may still have no idea what happens.
A follow-up could be:
Rewrite this for someone encountering the concept for the first time. State who does what, and under which conditions. Keep object names consistent. Preserve conditions that change the conclusion. Explain the behavior before naming the technical term. Borrow the clarity of ASD-STE100 for English without claiming compliance with the complete standard.
A more useful explanation would say: the system receives a request, performs the operation and saves a receipt. When the same request arrives again, it returns the receipt without repeating the operation. Assume the receipt still exists and the request contents have not changed.
That modest step supports everything that comes after it. A condition omitted from the prose will not automatically appear when the prose becomes a picture.
Second: give the question a place in a diagram
The post's own infographic demonstrates this. It places document structure, sentence rewrites, dictionary entries and writing limits in a single overview. You can compare a sentence with its rewrite, then follow the annotations to see why a word does not fit a particular use. Original infographic
A diagram gives relationships a location. You can point at a node and ask why a path passes through it, or point at a boundary and ask whether the conclusion holds there. A sequential explanation becomes something you can compare at once.
Tell the agent which relationship you need to see: causality, dependency, sequence or classification. Otherwise, it may produce a beautifully crowded overview.
Draw only the paths for the first request and its retry. Use three objects: request, operation and receipt. Distinguish executing the first request from returning an earlier result. Every arrow must have a defined meaning. Place the assumptions beside the diagram. Leave out other system components for now.
Review it with equally specific questions. Does each arrow represent the relationship you intended? Is its direction correct? Has an object silently changed its name? Without the surrounding paragraph, does the diagram still express the question?
We can learn from the source infographic's organization while remembering that an overview compresses detail. For a specific rule, return to the original documentation.
Third: turn “what if” into an action
Karpathy then suggests asking for HTML output. He is optimistic about models' frontend, interaction and animation capabilities. What interests me most is the ability to change a condition and inspect the model's response. Original post
Here is a small example. A warehouse has ten items. Request A reserves three, leaving seven. Its confirmation gets lost. You see a timeout and retry. Whether stock changes again depends on how the system handles that request.
Define a simple teaching model: with a saved receipt, retrying A returns the earlier result and leaves seven items. Remove the receipt before retrying, and this model executes again, leaving four. These are rules we chose for the explanation. Concurrent requests and real failure recovery sit outside it.
Ask the agent to make the question adjustable:
Build a single HTML file I can open locally. Show stock, request identity and receipt state. Add controls for sending A, retrying A and keeping the receipt. Let me predict before revealing a result. Include reset and the model's assumptions. Use sample data; controls change only this page's teaching state.
We have made a working version of this example. Keep the receipt or remove it. Send a new request B, or retain A's identity while changing its quantity. Those variations invite another question: what makes a request the same request?
This is what interaction adds. After hearing that retries do not repeat an operation, you can remove a condition and see whether the conclusion survives. Actual APIs have their own rules. Stripe's documentation, for example, spells out parameter and record-retention conditions for idempotent requests. Stripe documentation
Check rules and state when an agent builds this page. A responsive button establishes that an interface reacts. You still need to check whether each change follows the definition: seven items with a retained receipt, an explicit rejection for changed contents, a new operation for a new identity.

The same request under different conditions: seven above, four below. These are the teaching rules we defined.
Fourth: guide attention through the process
The format Karpathy is most optimistic about is a custom explainer video. His suggested experiments include a 3Blue1Brown-style explanation and ElevenLabs narration; he also mentions asking for free alternatives that use local compute. Original post
Video can resolve a different obstacle. A reader may see all the objects without knowing which change to watch first.
The revealing moment in our example is that the operation succeeded while the confirmation failed to arrive. Follow the request into the warehouse. Watch three items move into the reserved area and a receipt get saved. Follow the return message until it disappears. Pause. The viewer can now hold both states together: the sender thinks it failed; the recipient completed it.
Three-dimensional space and camera work have specific jobs. Stable locations help a viewer remember the request, stock and receipt. Following a message reveals sequence. Pulling back brings both sides into view. Labels should remain readable without the viewer having to chase them around the screen.
Give the agent that teaching job:
Make an explainer for this model. Show the difference between completed work and a missing confirmation, then explain the retry. Keep objects and identities consistent. Animate state or path changes. Include one prediction pause and the situations outside the model. Deliver narration, a storyboard and a representative sample before expanding to a complete film.
The storyboard, sample and pause are production choices I add to Karpathy's suggestion. They let us inspect the explanation while making it, before a polished film gives a mistaken account momentum.
Grant Sanderson's advice on the official 3Blue1Brown site emphasizes concrete examples and warns about unnecessarily moving equations and text. I like that constraint: every animation should correspond to a change the viewer needs to understand. 3Blue1Brown's advice
Why disposable software deserves care
Another important idea in the post is custom software that can be discarded. Building a webpage for one conversation or a video for one concept used to be hard to justify. Karpathy suggests that we should push what we ask for as creating these artifacts becomes more plausible. Original post
An understanding tool has a different purpose from a long-lived product. It can serve one meeting: expose a team's assumption, help a student explore a variable, or let the recipient of research locate the conditions behind a conclusion. When the question is answered, archive the page and keep the judgment and assumptions it helped clarify.
That expands what an individual can request. An agent can build a route into an unfamiliar field for this particular learner. When a problem feels distant, we can ask for something visible, something adjustable or a process we can follow.
In our OpenAI Dot project, replay let people point to a moment and discuss what happened. Karpathy's post makes me want to bring that ability into daily learning and work: receive a way into the result along with the result itself.

An AI-generated conceptual scene: a custom understanding tool can serve a single question.
Put it into the next agent assignment
You do not need all four formats. Tell the agent where understanding gets difficult, then choose a handoff that addresses it. Use prose to inspect conditions, a diagram to compare relationships, interaction to explore a variable, and video to follow a process. Keep a simple task's handoff simple.
I would add this to an assignment:
After completing the task, help me understand the result. What judgment will I need to make with it? Explain the conclusion, key relationships and assumptions. Choose a form that addresses my present question. If interaction helps, let me change one important condition. If animation helps, help me follow the process. Keep the sources and give me a counterexample or prediction question to check my understanding.
I have a job at receipt too: predict before operating, check the change against the rules, locate the conditions, and return to sources, code or run records when making a decision. An explanation may share the answer's omissions. A change in presentation does not automatically add independent evidence.
Try four outputs on a real question
Here is a question that needs no programming background: why send the James Webb Space Telescope about 1.5 million kilometres beyond Earth? This is an actual mission design choice. We use it as an explanation exercise; the page and storyboard below are ours.
NASA supplies two useful clues. Webb orbits the Sun near Sun–Earth L2, where the Sun, Earth and Moon can remain on the same side of its sunshield. Its faint infrared observations require a cold, stable environment; the shield helps separate the telescope from those sources of light and heat. NASA's orbit explanation, sunshield explanation
Writing gives the reason. Ask the agent for a short explanation: Webb needs to stay cold. The Sun, Earth and Moon bring light and heat. The location near L2 puts them on the same side, allowing the shield to sit between them and the telescope. Understand the design problem before memorising the term Lagrange point.
A diagram gives the positions. Place the three sources on one side, the shield in the middle, and the telescope on the other side. A reader can point at what must be shielded together. This is a relationship schematic; sizes and distances are not to scale.

Interaction lets us inspect a condition. In our four-format demonstration, display each source's schematic line, then hide the shield. Lines that ended at the shield can now reach the telescope. Does being far from Earth replace the shield? The page illustrates position and blocking; it does not calculate real temperatures.
Video gives the sequence. Begin with the telescope's need to stay cold. Introduce the three sources, add the shield, then pull back to explain the location near L2. Pause and ask the viewer to predict what hiding the shield changes. Attention has a route through the same facts.
Keep an accurate detail: Webb follows an orbit around L2 rather than sitting motionless at the point. Our schematic also does not portray it as hiding in Earth's shadow. NASA FAQ
An agent brief can specify sources and audience: use these NASA pages to explain the relationship between the shield and orbital location to a beginner. Deliver short prose, a relationship diagram, a page where a blocking condition can be inspected, then a video storyboard. Keep the facts consistent across forms and state the simplifications.
One real question now has a reason, visible positions, a condition to inspect and a guided sequence. We learn something concrete, and how to ask the agent for help understanding the next thing.
What excites me is the connection between agents' ability to create and people's need to understand. More questions can become visible, adjustable and approachable. As AI does more work, that handoff becomes worth designing with care.
