AI Builders Brief
?

Follow builders, not influencers.

2026.07.25

25+ builders tracked
BUILDER INSIGHTS
15
01
Boris Cherny Boris Cherny anthropicai

Boris Cherny

Opus 5 is a great model for coding, data analysis, design, biology, knowledge work.

More than any of these eval scores, what is most exciting to me is something else: Opus 5 is our least prompt injectable model yet. It is a bit buried in the system card, but across PI evals and red teaming, Opus 5 is very hard to prompt inject successfully.

And when layering defenses -- strong model alignment, combined with prompt injection probes, combined with Auto Mode in Claude Code -- the success rate for prompt injection attacks drops to ~0. This is new and exciting! More about this soon.

https://t.co/Tc7z2FqJhQ

X
02
Alex Albert Alex Albert AnthropicAI

Alex Albert

Just over 6 months later, Opus 5 now produces near-superhuman level spreadsheets and slide decks that match what a consultant would make. Things are changing fast. https://t.co/HKpFPP4hrM https://t.co/qv8OJoPXcf

Some of my favorite graphs from Opus 5 launch. We put a ton of work into making this model token efficient across domains while still raising the intelligence bar.

It feels very smooth to use and I prefer it over Fable 5 for many coding tasks. https://t.co/bScX0FgtLq

Welcome to the world, Opus 5. https://t.co/uLgA9FqQjR

X
03
Thibault Sottiaux Thibault Sottiaux OpenAI

Thibault Sottiaux

How to put this🙈ChatGPT Work is available globally for all paid plans across mobile/web/desktop. Probably already on your phone.

Puts a jetpack on your ChatGPT

Show me your pet https://t.co/WkrKZ8UyCj

X
04
Josh Woodward Josh Woodward VP, Google

Josh Woodward

Gemini does more than talk. It gets things done.

Today’s use case: Drop in a school calendar PDF. Tell Gemini to add every "No School" day to your Google Calendar. Done.

📍Gemini Spark is live now for all Google AI Pro subscribers in the US, expanding globally next.

X
05
Cat Wu Cat Wu anthropicai

Cat Wu

Claude Opus 5 is great at long-running autonomous work! Try it out and let us know what you think :) https://t.co/8LTbAOJJR0

X
06
Madhu Guru Madhu Guru CTO

Madhu Guru

There is a massive opportunity over the next few years for people who know how to take messy real-world workflows and adapt foundation models to them.

Doing that requires understanding how work gets done, designing evals, improving models through post-training, and building the feedback loops that continuously improve models.

That’s how a general-purpose model becomes exceptional for a specific domain.

Today, that skillset is still concentrated in a handful of labs.

Some people should be required to use AI for their writing.

X
07
Aaron Levie Aaron Levie CEO, box

Aaron Levie

Not sure if anything in tech has ever gotten as much broad-based support and alignment as this post and message.

The key now is that America should actually step up and continue to push open weights innovation. It’s good for the entire market, including the frontier labs.

Open weights accelerates the diffusion of AI across the economy by providing more options for specific customer requirements, can lower the cost of certain workloads, and you can tune models for areas of work that don’t always make sense in a horizontal frontier model.

Claude Opus 5 is out now, and it's a huge jump over Opus 4.8 on a wide number of capabilities.

At Box, we've been testing Claude Opus 5 with the Box AI Agent on Box's Complex Work Eval, our agentic benchmark that puts models through real enterprise document work end-to-end across a variety of industries. We saw meaningful gains in performance across some of the most complex enterprise tasks on unstructured data.

Here are a few examples of some of the wins we saw in testing:

* Due diligence (+17%): On a transaction due-diligence review, Opus 5 worked through the full set of required findings and flagged the ones Opus 4.8 missed. It stays thorough as the checklist grows instead of catching the obvious items and stopping.

* Life Sciences (+30%): On a target-identification task, it intersected several ranked datasets under a strict matching rule to find the targets common to all of them. Opus 4.8 over-included partial matches and dropped genuine ones.

* Legal (+12%): On a clause-by-clause contract review against a policy, Opus 5 scored every item and correctly cleared the clauses that were acceptable under an exception where Opus 4.8 both missed items and mis-scored the exception.

* Technology (+19%) and Healthcare (+13%) showed the same pattern: complete, precise, multi-step analysis over messy source material.

Overall, the reasoning, analytical, and data processing skills of Opus 5 outshine Opus 4.8 meaningfully. Going to be very powerful for enterprise agentic use-cases. You'll be able to build agents with Opus 5 in the Box AI Studio shortly.

Very happy to have Box sign this letter. We’re big believers in the power of open weights AI.

Open weights models help to drive the industry forward in a variety of ways to ensure more innovation, creativity, and diffusion of AI.

You get to have layers of the stack that emerge to post train models for highly specific purposes, which makes AI more useful in real world scenarios. Instead of waiting for just a few players to go deep in a domain, you get dozens or hundreds of attempts at that vertical, like in finance, life sciences, legal, healthcare, and more.

You get to see variance in how to handle safety and cyber risks. Instead of just one approach, you get a peek into what happens advanced capabilities can be used to build better systems to are used to defend systems.

You get alternative approaches to training and building AI models. In more compute constrained environments, you develop more novel and efficient approaches to model training, which every other lab can learn from.

And you get different cost structures for different workloads. High end and orchestration tasks can go to the closed frontier models and specific workhorse tasks can be done more cheaply.

Open vs. closed is not a zero sum battle. The reason why you want strong open weights models is because it pushes the entire AI industry forward.

X
08
Thariq Thariq anthropicai

Thariq

this is also available on the Claude Blog here: https://t.co/W84J5X640M

We removed ~80% of the Claude Code system prompt for our newest models, this is what we've learned about writing system prompts, skills and Claude.MDs for them. https://t.co/6DZwSrZjE9

Opus 5 rounds out our Claude 5 family beautifully.

I think it’s an incredible daily driver, pair it with Fable for planning, brainstorming or fixing the hardest bugs. https://t.co/S9GfhsS7v1

X
09
Garry Tan Garry Tan CEO, ycombinator

Garry Tan

"How fast a country adopted new technology over the last 200 years accounts for at least 25% of why some nations are rich and others aren’t today. The Ottomans got the printing press eventually but eventually cost them a lot." https://t.co/3Tg0oUz3dl

YC Startup School 2026 rehearsal at Chase Center

Welcome to San Francisco, future founders!

Tomorrow and this whole weekend are going to be 🔥 https://t.co/rwXOK3jBb0

To get macro productivity gains, managers and CEOs have to greenlight radically different staffing and workflow plans

To date I don’t think they’ve done it

Be prepared for this to take 10 years not 2 https://t.co/aNNYgDRp59

X
10
Amjad Masad Amjad Masad CEO, replit

Amjad Masad

It was always perplexing to me that VCs passed on Etched in the early rounds. Glad they woke up… and early believers get less dilution 🙂 https://t.co/wnVV5HVh8q

Will Anthropic sign?

If you work at Anthropic worth asking your leadership to sign or make their position clear: Are they for banning open weight models? https://t.co/99lCnlz5vW

If you haven’t used Replit for a while, you’re in for a big surprise. https://t.co/OrG0COztRW

X
11
Peter Yang Peter Yang

Peter Yang

Lying on the bed and talking to ChatGPT Voice to do work in codex is pretty great.

To do this well I think you have to remember the names of all your long running threads.

A beautiful day in Vancouver https://t.co/hz7D2PgNOz

I tend to agree with this.

Feels like pure software is really hard to monetize now for indie developers.

You need software + something else (e.g., services).

Would love to hear counter examples in the replies. https://t.co/8HMjINRd9y

X
12
Guillermo Rauch Guillermo Rauch CEO, vercel

Guillermo Rauch

Where's the ambition! https://t.co/kGiCsaFIPi

Ok let's see what you can do https://t.co/w3vt8GmwMK https://t.co/mn9q4sYkWa

Figma2React. It’s good https://t.co/ybjd12tIKG

X
13
Swyx Swyx smol_ai

Swyx

ok fine 4 new features - SmolForge is now getting customizable skins and spritesheet animations https://t.co/6uSOT102QS https://t.co/brUhUANaMU

the reason i am working on a new gsuite is because of extremely stupid defaults like this bullshit https://t.co/rO3dn0F26S

X
14
Matt Turck Matt Turck FirstMarkCap

Matt Turck

Under-discussed: the fundamental irony of the world's top AI researchers researching their way out of their own job, as they build recursive auto-research.

Big week in model routing:

Stripe rumored to acquire OpenRouter for $10B

Cursor Router announced on Wednesday

Runway Router announced yesterday

Also see routers at Databricks, Vercel, Cloudflare, Dataiku, AWS, Google (yes lots of differences behind the same term)

X
15
Nan Yu Nan Yu head of product, linear

Nan Yu

Endorse https://t.co/HxC3JFv6n9

X
BLOG UPDATES
1
PODCAST HIGHLIGHTS
1

STAY UPDATED

Daily builder insights, straight to your inbox.

Prefer RSS? Subscribe via RSS

ARCHIVE