Today’s lead · NVIDIA Technical Blog
NVIDIA’s AVO result says agent harnesses are now the battleground
NVIDIA says its AVO agent system lifted Claude Opus 5 to a 100.00 RHAE score on ARC-AGI-3’s public set. The useful takeaway is not “AGI solved”; it is that memory, supervision, tooling, and recovery loops can dominate model-only benchmarks on long-horizon tasks.

Top signals
4 moreTools & repos
6 selected
fx (by Vercel)
Vercel’s fx is a tiny open-source coding agent written in Zig and shipped as a roughly 6MB native binary. The v0.0.5 update adds Grok and Codex subscription support, project-local skills, security improvements, and bug fixes.

Antigravity IDE Extensions
Antigravity IDE Extensions put Google’s agentic coding platform inside VS Code, Visual Studio, JetBrains, and Zed, keeping agent conversations and shared context close to inline diffs, plans, debugging, and multi-step handoffs.

OneCLI
OneCLI pitches a self-hosted agent harness for teams, exposed in Slack and on the web. The useful angle is governance: sandboxed agents, policy controls, and avoiding direct access to real credentials.

Epho
Epho turns cloud coding agents into an API call: post a message, stream back work from Claude Code, Codex, or Opencode connected to a repo inside managed serverless sandboxes.

Supernova
Supernova connects startup data sources into Claude and Codex so teams can ask about revenue, pipeline, customers, usage, and operations without waiting on engineers or moving everything into BI first.

Plow Latch
Plow Latch is about letting agents operate a Mac while keeping data local and access scoped. That is the right problem area; the hard part is whether the scope boundaries are enforceable in practice.
Blogs worth your time
4 reads
LangChain’s trace judge is a reminder to fine-tune boring classifiers
LangChain and Fireworks fine-tuned Qwen-3.5-35B to detect “perceived error” in agent traces, reporting frontier-level accuracy at 10–100x lower serving cost. Useful pattern: specialize judges before paying frontier prices for every trace.

NVIDIA draws the security line below the agent harness
NVIDIA’s security post argues prompts and harness logic can steer agents, but runtimes and infrastructure must enforce authority. For builders, the clean takeaway is: agents propose; policy, identity, isolation, and audit live below them.

Google Research frames biomarker discovery as supervised agent work
Google’s Biomarker Discovery Framework uses multiple agents for hypothesis generation, statistical analysis, adversarial validation, and literature-grounded interpretation. The strongest note is restraint: it prioritizes candidate associations, not clinical validation or causal claims.

Ora is benchmarking whether the web is ready for agents
Vercel’s customer story is self-serving, but Ora’s benchmark pattern is useful: run multiple agent harnesses through real website journeys, trace where they fail, then use those traces to make products more agent-ready.
Community discussions
4 threadsAGENTS.md files are becoming repositories of agent scar tissue
A survey of top GitHub AGENTS.md files found lots of architecture, testing, commands, and very specific “don’t” rules. The best comment cuts through it: turn recurring prohibitions into linters where possible.
The MCP tool explosion is showing up as context and routing debt
Builders are debating whether agents should load all MCP tools, pre-scope them, or retrieve tools dynamically. The sharpest production warning: overlapping tools create hidden retries, and per-turn tool loading can wipe prompt-cache savings.
A honeypot for agents spending money without supervision
A Redditor built a disclosed “Certificate of Unsupervised Spend” tripwire to detect agents completing purchases without review. The thread quickly turns practical: liability, refund fees, and whether this is measurement or entrapment.
Claude users are fighting the model’s writing style, not just its code
Multiple Claude threads complain that recent Opus/Fable-style outputs are verbose, cryptic, or exhausting to review. The practical fix emerging from users is blunt output contracts: TL;DRs, reference docs, and constrained summaries.
Funding & acquisitions
3 moves
Starcloud adds $250M to its Series A for orbital AI data centers
Starcloud raised a $250 million Series A extension at a $2.3 billion valuation to build orbital AI inference spacecraft and secure launch capacity. The bet still leans heavily on Starship becoming frequent and cheap enough.

Nvidia takes a minority stake in data center developer Cloverleaf
Nvidia is moving further upstream in AI infrastructure by partnering with Cloverleaf, a data-center site and power-development middleman. Terms were not disclosed; reports cited by TechCrunch say Nvidia owns a minority stake.

Tross raises pre-seed funding for healthcare AI integrations
Tross raised pre-seed funding led by All In Capital, with DeVC participating. The startup is building APIs and integrations that let healthcare AI companies connect to EHRs, payer portals, and operational workflows.
Bengaluru radar
2 events
Ladies who ship 4.0
A sold-out women builders’ session in HSR: four hours to build with AI, automation, agents, or any stack and demo what shipped.

n8n Bangalore: Founders & Builders Mixer
Sold-out Koramangala mixer for founders and builders to discuss real business problems, AI automation approaches, n8n possibilities, and potential collaborators.

