Today’s lead · OpenAI
OpenAI launches a hosted Agents API backed by the Codex harness
OpenAI is turning more of the agent runtime into managed infrastructure: orchestration, sessions, context management, MCP tool connections, sandboxes, programmatic tool calls, compaction, and multi-agent delegation. Useful if you want long-running agents without owning the harness, but the execution environment and tool permissions still need real design.

Top signals
6 moreTools & repos
2 selected
AI Observability by OpenObserve
OpenObserve is aiming at the messy middle of agent ops: tracing cost, latency, quality, loops, failures, model calls, tools, services, datastores, and user sessions alongside normal logs, traces, and metrics.
AlexsJones/llmfit
llmfit is a Rust repo with a sharp utility promise: test hundreds of models and providers from one command to find what actually runs on your hardware.
Blogs worth your time
3 reads
Google’s ToolGrad flips tool-use data generation to answer-first
ToolGrad generates verified tool-use chains before writing prompts, then uses textual gradients to iteratively extend API workflows. Google reports higher pass rate, lower generation cost, and strong BFCL gains after fine-tuning Gemma-3 models.

NVIDIA shows what a tuned NIM serving stack buys on Nemotron 3 Ultra
NVIDIA’s Nemotron 3 Ultra NIM post is a concrete serving playbook: kernels, parallelism, prefix and state reuse, scheduler tuning, and MTP speculative decoding delivered 1,997 tokens/sec on 4xB200 at 50 TPS/user.

Credit Genie uses OpenWiki to make repo docs part of the code lifecycle
The useful pattern here is not another docs portal; it is docs as CI. Credit Genie runs OpenWiki nightly, opens update PRs from code changes, and points both engineers and coding agents at repo-local context.
Community discussions
5 threadsAgent builders are paying a schema tax on every tool turn
The thread nails a real agent cost problem: sending 40 verbose tool schemas every turn. The tension is recall versus token spend, plus whether pruning quietly hurts prefix caching and failure visibility.
Human handoff is an agent state-transfer problem, not a summary problem
Support teams are converging on a practical handoff rule: don’t dump a transcript into a queue. Persist the customer goal, actions tried, tool results, escalation reason, owner, and final outcome.
Inference margins look thin unless you own real optimization or differentiation
A LocalLLaMA thread frames inference as a brutal commodity business: customers push prices down while GPUs and electricity dominate costs. Commenters argue margins move to batching, topology, kernels, reliability, compliance, and customization.
Coding-agent memory needs versioned decisions, not an infinite chat dump
The thread is a useful reminder that memory for coding agents is curation, not storage. Builders want architectural decisions, conventions, and durable fixes retained without feeding every stale conversation back into context.
Anthropic’s AI labor scenarios spark debate over extreme assumptions and transition costs
The debate centers on Anthropic’s scenarios being explicitly non-predictive but still stark. Users focus on the extreme case’s zero-new-human-tasks assumption, labor share collapse, capital gains, and whether robotics makes “non-knowledge” absorption unrealistic.
Funding & acquisitions
1 moves
Graph AI raises $13.3M Series A for pharmacovigilance automation
Graph AI raised $13.3 million led by Insight Partners to expand in the US and Europe. The startup sells Graph Safety for adverse event intake, case processing, aggregate reporting, signal detection, risk management, and regulatory compliance.
Bengaluru radar
1 events
ElevenCreative hands-on workshop comes to Bengaluru
Free in-person ElevenCreative workshop with demos, a hands-on AI content challenge, networking, lunch, and project showcase in Bengaluru on September 18.



