When Thariq Shihipar joined Anthropic, he couldn't get his startup friends to try agentic coding, where an AI agent writes and changes code by itself. He recalls them saying their engineers didn't think it was good enough. Now, he says, less than 12 months later, it is "the default way that everyone codes." His own work has changed too. He no longer has to sell people on the tools. Now he teaches them to use the tools well.
Shihipar works on the Claude Code team at Anthropic. Claude Code is the company's coding agent: a program that lets Anthropic's Claude models read a codebase, run commands and make changes. He spoke to hosts swyx and Vibhu on Latent Space in an episode published on 29 September 2026. His main argument is that better agents have not made human skill less important. As the harnesses improved, he said, "the dominant problem now is, like, how do you use the agents." A harness is the software loop around a model that lets it use tools and carry out tasks. He called using agents a "high skill expression thing."
Prompting as writing for an audience
Shihipar calls prompting the "meta skill": the skill that improves all the others. He said many people think prompting no longer matters because they can type one sentence and Claude will do the work. He compared it instead to public speaking or writing "for a specific audience and that audience is Claude." What counts most, he said, is a mental model of Claude: what it does well, what it can "one-shot" (get right in a single attempt) and what it can't. Some expert users write short prompts that work because they understand both Claude and their own code so well. He said that makes their prompting look effortless even though the skill has a high ceiling.
The second skill is finding what he calls your "unknowns." As Claude takes on more kinds of work, he said, people are increasingly likely to attempt something outside their own expertise. The hardest gaps are "unknown unknowns": things you don't know exist.
He used design as an example. He is not a designer, so he asks Claude for "eight different mock-ups" and picks from them. A designer might instead supply reference sites, name a font and a look, list components to show, or connect a Figma board, and so give much more precise instructions. People without that vocabulary, he said, need to learn it, and Claude can help them. Game design shows the same problem. Plenty of people now try to vibe-code a game, meaning they build it by describing what they want and iterating on the generated code. They often find the result isn't fun. In a flying game, he said, a game designer might spend days on how the plane feels and responds to the controls.
The discussion turned to "taste." Shihipar said he is torn on the word because it can sound elitist, as if only founders have it and engineers don't. He argued that engineers have plenty of taste for their own kinds of problems. He quoted Jason Liu: to have taste, "you have to eat." In Shihipar's version, that means iterating, working out what you like and building a domain vocabulary, then bringing all of that together when you write a prompt. The conversation also noted that sometimes people only find out what they want when a model produces it and it simply feels better.
These ideas repeat his July 2026 field guide to Anthropic's Fable 5 model. It separates explicit requirements, known gaps in knowledge, unstated preferences and concepts a user has never met. It suggests learning conversations, architecture interviews and competing prototypes, because a prototype makes unspoken preferences concrete enough to discuss.
Claude Code's Ask User Question tool builds the same idea into the product. Shihipar has a background in human-computer interaction and described the tool as the first time the model was good at "elicitation": asking questions to draw out requirements. Some users just want the agent to do the work. He believes, though, that "pretty much everyone" knows less about their problem than they think. Details such as the data schema are best settled before implementation starts. His earlier account of Claude Code's tool design says a dedicated question tool worked better than earlier attempts: it offers structured choices and pauses until the user answers.
Voice, long prompts and the cost of redoing work
The conversation compared two very different habits. One was rambling into a voice-dictation tool for a couple of minutes and hoping the model sorts it out. The other was spending a solid 30 minutes writing a first prompt for Fable, because models now run for longer and are hard to redirect once they start. Both habits came from the hosts' own experience. Shihipar said voice prompts are not necessarily worse. What matters is how much information the prompt contains, not its format. A model can follow a speaker who changes their mind halfway through. Many people find talking easier than typing, he said, and if that gets more information out of them, it is better.
The practical payoff is less rework. Shihipar said people often hit their usage rate limits in the same way. The agent does a lot of work, they don't like it, and they ask it to undo and redo. Then they go round in circles: "nope, don't like that design," try this, you messed that up. Often, he said, the model could have got it right the first time with more upfront time or better context. Those loops use up much more of a user's allowance.
The useful context goes beyond the goal itself. He suggested telling the model whether you are building a prototype or production software, and where it should or shouldn't spend compute. The model, he said, doesn't intuitively know how much you want to spend on a task, so sometimes you have to give it permission or tell it not to.
Matching effort to the task
One way to express that is through Claude's effort setting, which controls how much work the model does on a request. When a host asked how to tell whether they were using too much effort, Shihipar gave a rough split by type of work:
- Low for UI work and fast iteration.
- Medium for building something like an API, where you want enough edge cases covered.
- High or max for code review and security.
He said higher effort changes security evaluation results a lot, but ordinary software engineering not as much. In software work, the extra effort mostly goes into verification and edge-case testing. Knowing how these settings behave across different kinds of work, he said, is "part of the job."
A host asked whether this was intuition or based on evaluations. He pointed to his analysis of Terminal-Bench tasks, a set of benchmark problems for coding agents. His published effort guide gives the details. It compares 70 shared Terminal-Bench 3 tasks, leaving out four GPU tasks, with five attempts per task. It finds that extra effort often buys verification, edge-case handling and independent judgment. On one HTML-sanitizer task, Fable 5.1 went from one success in five attempts at low effort to five in five at extra-high effort, adding deeper parser checks and adversarial testing. The guide also finds limits. More effort did not reliably rescue a wrong approach. The runs also used internal configurations, including disabled Fable safety safeguards and restricted internet access for security tasks, so the scores are not directly comparable with public leaderboard runs. The guide's advice is similar to his podcast summary: low effort for fast iteration, medium for ordinary feature work, and higher settings when verification or difficult independent reasoning matters.
Shihipar also made a forecast about which model to choose. He said it is "not quite true yet, but it's very close" that the frontier models (the most capable ones available) will beat smaller models on both capability and token use for almost everything. His reasoning is about verification. A perfect model would do the work once and not need to check it. He said he already finds himself telling Fable it doesn't need to launch the Chromium browser and screenshot everything to prove a change worked. As models get smarter, he expects them to spend fewer tokens checking simple work, making them more efficient than smaller models. Tokens are the units of text models process and bill by. A task-cost guide by Addy Osmani shows how this could happen. A task's cost includes retries and a growing conversation that is re-sent on each turn, so a model with a higher price per token can sometimes cost less per finished task. The same guide notes that switching models mid-session starts with an empty prompt cache, which loses the savings from reusing earlier context.
Asking the agent what it decided not to do
The Terminal-Bench transcripts led to another of Shihipar's tips: ask the model to write decision notes or implementation notes as it works. In almost every task, he said, the model considered the correct solution and then decided against it. At higher effort levels, that pattern accounted for most of the failures. It is "very rare," he said, that the model simply doesn't know how to do something. With notes, you can review the choices and tell the agent to do the thing it passed over.
He said Fable 5.1 already points out more of its decisions in its output, but making this explicit in the harness works better. His field guide proposes a temporary implementation-notes file that records deviations from the plan, edge cases and conservative choices. The reviewer then gets a record of decisions alongside the finished code.
Keeping instruction files lean
The conversation then turned to instructions that last longer than a single prompt, such as project goals kept in CLAUDE.md, a file of standing instructions for Claude. A host noted Shihipar's documented dislike of AGENTS.md, a similar file used by other coding tools. He said Anthropic will support it anyway, because maintaining separate files for different tools is such a pain. He went further on CLAUDE.md itself: "in the limit, Claude.md goes away," he said, and it may already be better to start a new project without one.
His concern is old workarounds. If a failure keeps happening, he suggested, add it to the file. But failure modes change from model to model, even between Fable 5 and Fable 5.1. A running list of old failure modes will "probably over-constrain Claude." He said Anthropic has added a way to evaluate whether a skill, meaning a reusable package of agent instructions, actually improves results, while admitting that this costs tokens and is not perfect. Anthropic's documentation on project memory fits this approach. It says CLAUDE.md is loaded as context rather than enforced like configuration, recommends short, specific and consistent instructions, and suggests keeping the file under 200 lines.
The discussion also brought in a writing framework from outside AI. Clear prompting, the argument went, looks a lot like executive communication. Barbara Minto's Situation–Complication–Question–Answer structure, taught in a workshop Heavybit described in 2019, sets out shared context, what changed and the decision to be made. When you don't have the answer yet, you can still write the first three parts.
Understanding the result before shipping
Shihipar's advice continues after the code is written. His field guide suggests pairing each implementation with an HTML explanation and a quiz before the change is merged. The point is to check that the person understands the change, rather than assuming that working code means they do. He also mentioned a short explain-it-like-I'm-five skill that can be installed as a plugin. He said its key instruction is simply to give the big picture in few words. It came from people at Anthropic trying to follow very complicated incidents, and he described it as "shockingly good" at cutting through the noise. Artifacts, the documents and interactive pages Claude produces, often contain too much text, he said, and people don't read them.
Quizzes were harder to sell. The conversation described a test-your-understanding approach: offer a few multiple-choice answers, and a wrong answer shows a gap between what you think is happening and what is really happening. It works like Ask User Question in reverse, asked after the work instead of before it. Shihipar said this is something everyone loves to talk about and very few people do, because most people don't want to be quizzed. The discussion closed the topic with a rule against sending colleagues AI output you haven't understood: "before you send stuff, you should at least know what's implemented."