Today’s lead · Ars Technica
Meta’s Muse shows the security cost of privileged desktop agents
Muse is a useful warning for anyone shipping agentic desktop software: broad OS and account permissions turn prompt-level mistakes into real compromise. Ars reports a zero-day that lets local apps and terminal commands take over the macOS agent, undercutting Meta’s privacy-and-security claims.

Top signals
6 moreTools & repos
4 selected
Arcjet
Runtime security primitives for AI apps: prompt-injection detection, tool-call authorization, sensitive-data redaction, and bot or abuse blocking before an action executes.

Superset Mobile
An iPhone control surface for coding agents: resume desktop workspaces, review diffs, and merge PRs while Claude Code, Codex, and other terminal agents do the work.
akitaonrails/ai-memory
A Rust repository for long-term memory across agent coding CLIs, aimed at preserving context and easing handoff between different agent vendors.

Hyrax AI
A codebase-wide improvement agent that builds architectural context, prioritizes issues across engineering domains, writes fixes, verifies against tests and checks, and opens GitHub PRs.
Blogs worth your time
3 reads
Agent evals should grade the world state, not just the tool call
NVIDIA’s useful framing: tool-call accuracy is only a component metric. Production agent evals need executable environments, final-state checks, consistency ranges, step traces, and cost per successful task.

Block-pruning LLMs as an Ising optimization problem
Multiverse Computing reframes depth pruning as constrained binary optimization, using Hessian-derived block interactions instead of independent importance scores. The reported gains matter most under aggressive compression.

Alexander Mattick on why information is expensive in intelligence
A dense MLST conversation on inference, energy-based models, flow matching, RL constraints, world models, and why priors and constraints are not optional when information is costly.
Community discussions
4 threadsWere AI sandbox “escapes” really just bad network isolation?
The useful tension: true air gaps make many agent tests unrealistic, but calling software barriers “air-gapped” muddies the lesson. Builders need explicit threat models, egress rules, and route analysis.
A practical pattern for multi-day, multi-machine coding agents
The workflow is a useful counter to “agent = chat loop”: write a short design doc, use a reasoning-heavy coordinator, fan work across machines, and monitor workers over days.

Local agents get a serious workstation pitch with the M5 Ultra Mac Studio
The post argues local AI is becoming a workflow choice, not just a privacy hobby. The tension remains cost: expensive hardware competes with years of cloud subscriptions.
Should Claude Code spend tokens on git, builds, and deploys?
The practical split: let Claude drive repetitive CLI workflows, but make it call scripts and watch exit codes when possible. That preserves tokens and makes automation reusable.
Funding & acquisitions
2 moves
Corridor raises $25M seed for an AI health benefits brokerage
Corridor is betting AI agents can make small-business health benefits brokerage economically viable. Humans advise customers while agents handle admin work like network checks, scheduling care, and insurance-info updates.
Kos raises $12M for AI agents in data center finance
Kos is narrowing agentic finance to capex-heavy operators: data centers, energy, and manufacturing. The claim is audit-grade workflow execution across invoices, pay apps, POs, ERPs, chat, calls, and shared files.
Bengaluru radar
1 events
AI@Adyen: Transforming the Fintech Landscape
Today at Adyen Bengaluru: AI in fintech engineering, a live Adyen payment integration demo, leadership keynote, dinner, and networking. Sold out; bring photo ID if confirmed.





