Today’s lead · OpenAI Developers
OpenAI says Codex has 7M+ weekly users and ships GPT-5.6, parallel work, inline edits, and PR workflows
OpenAI says Codex now has more than 7M weekly users and has received 150+ updates in two months, including GPT-5.6 and Ultra, parallel work with /goal, faster computer use, AppShots, inline edits, Sites, mobile and SSH workflows, and PR review-to-merge flows. The update positions Codex less as a coding assistant and more as an end-to-end software delivery surface.
Top signals
9 moreTools & repos
10 selected
Bonsai 27B
PrismML released Bonsai 27B, an Apache-2.0 multimodal model based on Qwen3.6 27B, in low-bit variants sized for local deployment: a 5.9 GB ternary build at 1.71 effective bits per weight and a 3.9 GB 1-bit build for phone-class footprint. Together AI says its ternary API build has 262K context, vision input, and retains 95% of full-precision quality.
MOSS-VL-Realtime
MOSS-VL-Realtime is an open-source 11B vision-language model for continuous video streams under Apache-2.0. It can answer questions while still watching, revise or interrupt responses as scenes change, stay silent when more evidence is needed, and supports a 256K-token context window with Chinese and English multimodal understanding.

Unsloth Gemma 4 NVFP4 quants
Unsloth released Dynamic NVFP4 quantized Gemma 4 models for NVIDIA Blackwell GPUs, claiming 1.5× faster local inference. Gemma-4-12B NVFP4 fits in 11GB VRAM, while the 26B-A4B model reaches 13K tokens per second on B200, using Unsloth’s mixed FP4/FP8/BF16 quantization strategy.
WANDR
Perplexity open-sourced WANDR, an internal benchmark used to build deep and wide research capabilities inside Perplexity Computer. It is designed for high-volume, evidence-heavy research agents that must search broadly and deeply rather than answer from a narrow retrieval path.
OpenRouter MCP
OpenRouter shipped two weeks of MCP updates that let an agent discover, rank, test, and report on models without leaving the editor. For teams comparing model quality, latency, and price, this turns model selection into an agent-callable workflow rather than a separate dashboard task.
Sailboxes
Sailboxes are cloud machines with persistent state, auto-sleep, and pricing from $0.015 per active vCPU-hour, built for long-horizon AI agents that need to resume work over time. Sail says the underlying architecture live-migrates VMs based on actual resource usage and was demonstrated by repeatedly migrating a Minecraft server every two minutes without player-visible disruption.
Mem0 Anthropic connector
Mem0 is now an official Anthropic connector available from Claude’s Connectors directory, with no CLI, config file, or MCP setup command required. Once connected, Claude on the web, desktop, or Cowork can remember context across weeks rather than only one session.
screenpipe
screenpipe is an open-source, local-first tool that records and learns how you work, then turns that activity into searchable memory, SOPs, and AI agents. The project claims 20K+ GitHub stars, 1,900+ forks, and 130+ contributors, making it one of the more mature personal-workflow capture layers for agents.

PgDog
PgDog is positioned as a way to scale PostgreSQL without changing application code. For teams hitting database limits while moving AI features into production, the practical value is reducing app-layer migration work around Postgres scaling.

LACT
LACT controls AMD, NVIDIA, and Intel GPUs on Linux with monitoring, overclocking, fan curve management, power configuration, settings profiles, and OpenTelemetry metrics export. Because configuration is handled by a system service, it can also be useful for headless GPU boxes used for local model work.
Blogs worth your time
8 reads
How to debug coding agents with LangSmith traces
LangChain shows how LangSmith traces expose model calls, tool calls, shell commands, MCP activity, subagent fanouts, retries, costs, and errors across Claude Code, Codex, Cursor, GitHub Copilot Chat, OpenCode, and related agents. The useful lesson is that coding-agent failures are often handoff, context, or tool-selection bugs that only become obvious when the intermediate trace is visible.

Lessons from 5,000+ Kagglers on improving AI reasoning
NVIDIA distills the Nemotron Model Reasoning Challenge into practical patterns: verify intermediate reasoning, compress traces to fit token budgets, separate reusable knowledge from new problem solving, use tools to generate and audit training data, and evaluate by task type rather than only leaderboard position. It is a useful field report for teams trying to improve reasoning without changing the base model.

How to run an autoresearch workflow with RL agent skills and NVIDIA NeMo
NVIDIA walks through a skill-based autoresearch workflow where a coding agent sets up environments, launches experiments, monitors metrics, and iterates on reinforcement-learning tasks using NeMo RL and NeMo Gym. The concrete example moves Qwen3-VL-2B from 25% to 96.9% accuracy on a custom task, making the post useful for teams exploring agents as ML experiment operators rather than code generators only.

Post-train NVIDIA Cosmos 3 in one day using agent skills
NVIDIA shows how TAO agent skills, LoRA, and AutoML can automate Cosmos 3 Nano post-training for video question answering. The reported Woven Traffic Safety result improves exact-match accuracy from 54.41% zero-shot to 93.35% in under a day, with deployment through Cosmos 3 Reasoner NIM as OpenAI-compatible endpoints.
Scaling PyTorch-authored generative AI inference across multiple GPUs with TensorRT 11
The NVIDIA Developer post highlighted by PyTorch explains how PyTorch generative pipelines can be converted through Torch-TensorRT into TensorRT engines for C++ production deployment. It covers TensorRT 11.0 multi-device inference, NCCL-backed distributed collectives, context parallelism, and tradeoffs among AllGather KV, Ring Attention, and DeepSpeed Ulysses for long-sequence media generation.
How Claude performs on robotics tasks
Anthropic’s robotics evaluation tests language models across classic control, simulated quadrupeds and humanoids, a robotic arm, and a real Unitree Go2 with different levels of control abstraction. The practical takeaway is architectural: models mostly fail when directly driving joints, but perform better when supervising pretrained controllers or using higher-level control interfaces.

Xiaomi-Robotics-U0: a 38B world foundation model for embodied synthesis
The paper introduces Xiaomi-Robotics-U0, a 38B multimodal autoregressive model that keeps general image and video generation in the training mix while learning multi-view embodied scene, transfer, and video synthesis. The reported synthetic-data result is notable for robotics teams: out-of-distribution success for a real-robot policy improves from 36.9% to 63.2%.
How to manage AI investments in the agentic era
OpenAI frames enterprise AI investment around useful work per dollar, efficiency gains, and scaling high-value workflows. It is a useful operator lens for teams moving from demo spend to agentic systems that need measurable economic output.
Funding & acquisitions
8 moves
Reflection AI signs $1B compute deal with Nebius
Reflection AI, a U.S. open-model startup founded in 2024 by former Google DeepMind researchers, signed a $1B compute deal with Nebius for access to Nvidia’s latest chips. The agreement follows a similar SpaceX compute arrangement and shows how open-weight labs are locking down infrastructure as a core strategic asset.

DeepSeek reportedly seeks $1.5B at a $71B valuation ahead of a possible IPO
DeepSeek is reportedly preparing for a 2027 IPO and is in talks to raise about $1.5B at roughly a $71B valuation, after raising $7B a month ago at about $50B. TechCrunch reports its cloud service runs on Huawei chips and that it accounted for nearly 23% of enterprise-focused AI gateway Vercel’s token volume in June.
Valarian raises $50M Series A led by NEA for sovereign AI infrastructure
London-based Valarian closed a $50M Series A led by NEA to build AI infrastructure where each model, agent, and workload is sealed in its own enclave, runs on customer-controlled infrastructure, and uses keys held by the customer. The pitch targets enterprises worried about IP exposure in third-party AI stacks.

Meta invests in CRED through a Series H primary round and broader $900M transaction
Entrackr reports Meta’s $900M investment marks CRED’s first fundraise in nearly four years, valuing the fintech at about $4.53B post-money. Filings show Facebook Overseas Inc. will invest ₹5,107 Cr, or about $537M, through Series H CCPS for an 11.84% primary stake, while CRED also expanded its ESOP pool by more than $110M.

E3 Electric.Ai raises ₹100 Cr Series A for AI-powered electric scooters
Bengaluru-based E3 Electric.Ai raised ₹100 Cr in a mix of equity and debt led by BluVenture Holdings ahead of the commercial launch of its E3 TRION scooter. The company says its AI-powered two-wheelers use AI-enabled safety systems, predictive diagnostics, connected vehicle technologies, modular powertrain architecture, and battery intelligence.

Udaan lines up $160M structured financing to repair its balance sheet
B2B ecommerce company Udaan announced a proposed $160M financing transaction comprising fresh equity, new debt, and debt-to-equity conversion. The transaction includes a roughly $45M private credit commitment and comes after creditors initiated proceedings against Udaan’s Singapore parent over defaulted convertible notes.

Hero MotoCorp approves up to $104M investment in Ather Energy
Hero MotoCorp approved an investment of up to ₹10B, or about $103.95M, in Bengaluru EV maker Ather Energy through a preferential allotment of shares or convertible securities. Hero held about 29.48% of Ather as of June 30, with the final stake dependent on the pricing and structure of the securities issue.

Hisabkitab raises seed funding to build AI agents for SME accounting
Surat-based Hisabkitab raised an undisclosed seed round from angel investors and HNIs at a ₹20 Cr valuation. The AI-powered, cloud-native accounting platform plans to build an AI Intelligence Layer with Audit, Tax Preparation, Accounts Receivable, and Accounts Payable agents for small and medium businesses.
Bengaluru radar
1 eventsThe Apache Iceberg Edition
Data & AI Forum is hosting an in-person Bengaluru meetup on building an Apache Iceberg lakehouse and querying it with multiple engines, followed by a technical deep dive and networking. The announced speakers include Shubham Baldava, CTO at Datazip, and Avinash Upadhyaya, Platform Engineer at Platformatory.








