AI Builders Brief
?
← BACK TO TODAY

Follow builders, not influencers.

2026.09.11

25+ builders tracked
BUILDER INSIGHTS
15
01
Thibault Sottiaux Thibault Sottiaux OpenAI

Thibault Sottiaux

Scaled agents on demand.

This is pretty much the infrastructure that runs under the hood for ChatGPT Work, all wrapped up in an API which you can use to get started in < 1 min. Happy building. https://t.co/GhmuBkfjP1

Do you pronounce it data or data? This is the way everyone makes dashboards and learns about the business at OpenAI. Can't live without it. https://t.co/9NAremYs1h

To make sure our current users have an incredible experience and continued access to Astra, we are going to pause subscriptions to our $200 Pro plan. These put the most strain on our systems and we wanted to take the smallest step that allows us to continue giving the broadest access possible. All other plans and the api remain available.

There is no impact to existing accounts and we are working on adding more capacity as fast as we can. Thanks!

X
02
Thariq Thariq anthropicai

Thariq

Try this prompt in Claude chat to give it more context about yourself:

Interview me in depth using free text, or askuserquestion tool when multiple choice works, about relevant parts of my life you don’t know about yet and save it all to memory. https://t.co/VnfnFo7h4q

X
03
Josh Woodward Josh Woodward VP, Google

Josh Woodward

Gemini, now on Windows!

Get it: https://t.co/mKpROGeCl3 https://t.co/59jHumlY0X

X
04
Google Labs Google Labs

Google Labs

It’s time to ☁️ dream bigger ☁️

Dreambeans is officially available to all US users (18+) on iOS and Android, free of charge. No subscription required.

You can also now connect @Geminiapp to Dreambeans. Dreambeans will build off the nuance and understanding from your chats to surface even more insightful and personalized daily stories.

Get your daily, freshly brewed collection of stories here: https://t.co/jCdAzMHvOD

X
05
Boris Cherny Boris Cherny anthropicai

Boris Cherny

The latest Threat Intelligence report is an absolutely terrifying and important read.

As models become more intelligent, without the right safeguards and monitoring they also become more dangerous. Many capabilities are dual use: a model that codes well can be used to hack critical infrastructure; a model that assists with biology research can also be used to engineer the next pandemic.

These issues are complex, thorny, and increasingly important for everyone to understand so that the world can weigh in and respond to rapidly escalating risks.

https://t.co/0rnV2HjPP6

Hey ████,

I think there is room for both.

1. Prototypes and other throw-away code can be treated as totally black box. If you’re going to throw it away anyway, and if the blast radius of it breaking is low, it doesn’t need to be perfect.
2. Production code written by Claude should have a higher bar than if it was written by a human. At Anthropic, we have many guardrails in place to make sure this is happening: lots of lint rules, lots of tests, Claude-driven end to end tests, Claude-powered fuzzers running daily, automated code reviews and security reviews, automated code refactoring, and so on. Without these, you can end up with a mess that is hard to maintain down the line. Luckily, the model makes it increasingly easy to do these well — run a few daily routines, use Claude Code Review, etc.

Your job is to hold the bar on code quality. If Claude’s code doesn’t meet the bar, try:

  • Using the latest frontier model (Opus 5 or Fable 5.1)
  • Increase effort to high or xhigh
  • Invest in your CLAUDE.md and skills to succinctly teach Claude how to work in your codebase

If all else fails, steer Claude more when you work with it, or have Claude fix accumulated debt and rewrite your codebase to make it easier to work with. Or, wait for the next model.

Best,
Boris

Every day, I get a lot of of emails and messages like this one. I try to respond to as many as I can.

Sharing my response below, for anyone else in a similar situation. What do you think? https://t.co/z1GtgK14RM

X
06
Peter Yang Peter Yang

Peter Yang

In my humble opinion, for getting shit done Sol > Astra

X
07
Aaron Levie Aaron Levie CEO, box

Aaron Levie

Some more tales from the road. Met with a couple dozen technology leaders this week across banking, media, information services, insurance, and consulting to discuss agents in the enterprise.

Some of the biggest trends right now:

* Cyber! Everyone nervous about the growing rate of vulnerabilities coming at them from AI, and the implications of the OpenAI Hugging Face incident. The conversation is not as existential as it is in Silicon Valley, but still highly concerned and pragmatic about what to do about it operationally in their environments. Lots of new discoveries due to AI, and still hard to keep up with all the changes they have to execute now.

* Model battles persist. Most companies are deploying multiple frontier models within their enterprise. Too hard to standardize on anything and seeing different preferences across their teams and use cases. But the dollars are still concentrated on just a few vendors. Open weights still in infancy at scale in most of these organizations, often due to lack of domestic “frontier” OSS options. Plenty of appetite for more options here, but so far few places to go.

* Agent security and identity. Somewhat tied to Hugging Face, there’s much more awareness to the new challenges around agent security and identity management in a world when agents are trying to get into every system they can. In a perfect world enterprises could setup identities for all their agents and control what they’re doing, but of course sometimes the agent needs to act exactly as the user as well.

* Process reengineering. Most companies realizing that the big upside of agents is when they can change the actual workflow itself to get the full gains from AI. Far more ROI when companies can adjust their workflows to support agents changing how the work happens instead of just layering on agents into the existing flow. But the big question is who can actually tackle driving these changes, where does that live, etc. Best lessons were still around embedded FDEs in the functions.

* Ruthless adjusting of architectures. Most companies had examples of changing systems out multiple times just in the past year or two with different vendors. I probably haven’t heard “we tried X and it didn’t work so have gone with Y” more than in today’s environment. The lesson here is that because innovation is happening so fast, no one hangs around until a vendor gets something right, they just move on to the next one.

* Evals! Still very early for most companies to have a good grasp of evals of their workflows. A few customers out of a couple dozen called this out - huge opportunity right now for enterprises to have a good sense of how their work actually happens and how well AI is doing against it.

* Legacy systems still a hurdle. As always, legacy systems still remain a mainstay issue that holds back enterprises from rapid adoption of AI in enterprises. Data is fragmented across legacy environments that weren’t built for an agentic world. Companies spending a lot of time just cleaning up these old platforms.

Many more topics, but these tend to be some of the more top of mind items at the moment in the enterprise.

Excited to partner more deeply with OpenAI so you can work with your enterprise content from Box securely in ChatGPT.

Software continues to go headless, and we’re massive believers that the future of work will be AI agents that can process data and execute workflows anywhere. https://t.co/CNZv5CH8bn

X
08
Peter Steinberger Peter Steinberger OpenClaw

Peter Steinberger

Trading tokens for an AC. Who knew SF could be so hot 🫠

This makes a lot of sense. Duplicating logic is no longer painful. Abstractions still are. https://t.co/q4jqVBW00r

Better hop on soon. Astra demand is growing too fast! https://t.co/MgCFw04sxJ

X
09
Madhu Guru Madhu Guru CTO

Madhu Guru

How to build great evals - part 10

Measure the steps, not just the result.

Much like high school math, it isn’t sufficient just to get the right answer, the steps to get there are critical.

Two agent trajectories might produce the same answer (42!!).

but one of them searches the right sources, retrieves the right document, makes 4 clean tool calls and calculates the result. The other makes 17 calls, searches the same thing 3 times, recovers from 2 errors and eventually gets there.

It’s clear which one is better.

Here’s what you need to do:
1/ clearly define your whole workflow
2/ define the tasks in each step
3/ think through how you measure each step - separate evals or is it a slice of a bigger eval
4/ define your median and hard tasks - reflect them in your evals

Now any time you look at eval results, study the steps first and the final results next.

X
10
Guillermo Rauch Guillermo Rauch CEO, vercel

Guillermo Rauch

A computer for every agent, in every region https://t.co/3Bzea8XIY0

Each day ~10 million deployments are made on Vercel, with 2.35 billion made to date.

Vercel is one of the most heavily multi-tenant systems in the world. Billions of application deployments co-exist and are accessible ('routable') on our CDN at any given moment.

Underlying our CDN is a global metadata store that synchronizes within hundreds of milliseconds, globally. e.g: when you roll back, change config, add routes, etc.

We just made this system 91% faster at p99, and in the process sped up the build→deploy pipeline. All of this while the system is under immense pressure from the growth in agentic deployments.

Great read on the internals of Vercel from our CDN engineering team:

We made deployments faster again https://t.co/lrtQdx1iK1

X
11
Amjad Masad Amjad Masad CEO, replit

Amjad Masad

Chat with the etn bros! https://t.co/lMH3Pqbeat

Chat with PG in London! https://t.co/E0Gs90oQRc

Lots of risk with AI. I worry a lot about cybersecurity for example. However, “extinction risk” — literally 100% of humans die — is not remotely one of them. https://t.co/48n1fYmyh7

X
12
Zara Zhang Zara Zhang

Zara Zhang

Why is computer use still so painfully slow??

X
13
Nan Yu Nan Yu head of product, openai

Nan Yu

Yet normies use Google and Instagram and Zillow and Doordash all day every day.

Still. Early. https://t.co/YO92sTMaqd

Don't call it a private equity acquisition for a huge discount from peak valuation.
Call it an Italian goodbye.

X
14
Matt Turck Matt Turck FirstMarkCap

Matt Turck

This conversation with @RichardSocher is also available on Spotify, Apple Podcasts and here on YouTube (like and subscribe!):

https://t.co/8gsssP8fcu

When AI builds AI: my conversation with @RichardSocher about RSI and scientific progress.

00:00 Intro: AI That Improves Itself
00:55 Why Scientific Progress Is Slowing
03:08 The Labyrinth of Human Knowledge
05:59 Can AI Put Science Back Together?
07:57 How LLMs Learn Biology and Proteins
10:56 Next-Token Prediction as a World Model
16:44 Can AI Generate Truly Original Ideas?
17:32 Simulations, Verifiers and Superhuman AI
22:18 The Path to Recursive Self-Improvement
24:49 Why AI Hallucinations Can Drive Discovery
27:42 From Reading Biology to Writing It
31:31 Can AI Accelerate Drug Discovery?
33:31 Will AI Help Cure Cancer?
38:03 AI Breakthroughs in Biology, Energy and Materials
40:19 Will Some Societies Reject AI?
45:07 Building the AI Economist
52:22 The Scientific Data Bottleneck
53:41 The Four Pillars of the Eureka Machine
55:01 Teaching AI the Rules of Reality
57:44 Simulations and Virtual Cells
1:00:40 Self-Driving Robotic Laboratories
1:02:51 Agent Swarms and Open-Ended Discovery
1:04:30 The Compute Bottleneck
1:05:44 Inside @Recursive_SI
1:07:33 What Recursive Will Build First
1:10:10 How Do We Define Intelligence?
1:11:32 How Far Can Intelligence Go?

X
15
Nikunj Kothari Nikunj Kothari Partner, fpvventures

Nikunj Kothari

This post is probably my personal record from thought -> publishing.. lacks some of the usual polish!

A little behind the scenes:

> voice memo while driving to work
> more voice memos throughout the day
> 30 minutes of furious writing between two meetings
> quick read through & publish

@claudeai please fix your voice transcription 🙏

https://t.co/pPmfsH024O

Three truths in early stage venture right now:

1) everyone wants to raise a $50 million seed
2) everyone thinks they will hit $30 million ARR next year
3) every hot tranched seed round magically ends up at the ~$300 million valuation

X
PODCAST HIGHLIGHTS
1

STAY UPDATED

Daily builder insights, straight to your inbox.

Prefer RSS? Subscribe via RSS

ARCHIVE