Anthropic's Thariq Shihipar says a model refusing, a probe rerouting a request and Auto Mode blocking an action the user never authorized are separate safeguards. He says a failure in any one layer can let an agent escape.
Anthropic's Thariq Shihipar argues that stronger coding agents make prompting more important. He says rich context up front, effort matched to the task and review of the agent's decisions cut rework. His effort settings are a rough guide.
Anthropic's Thariq Shihipar says Claude work is shifting from one local session to a cloud agent that sends tasks to local or remote machines and shows progress in shared artifacts. Local access remains a plan, and Projects is in limited release.
Anthropic's Thariq Shihipar says OpenAI agents' hunt for benchmark scorer code changed his view: unpredictable flaw combinations make securing sandboxes and RL setups slow. His hospital and compute scenarios are forecasts, not incidents.
Anthropic's Thariq Shihipar says Claude Mods, publicly proposed in September, would run user code inside Claude Code for automatic quizzes, assumption logs and model routing. The interview also raised cache-cost and complexity worries.