DeepSeek Harness v0.1: The Stack Comes Apart
DeepSeek’s developer-preview agent harness treats models, tools, sessions, sandboxes, orchestration and UI as replaceable plugins—but practical interoperability remains unproven.
THE LOG — TRANSMISSIONS FROM ALTITUDE
What we learn running AI systems every day — workflows, growth loops, tools, and the occasional hard lesson. Written down so you can reuse it.
DeepSeek’s developer-preview agent harness treats models, tools, sessions, sandboxes, orchestration and UI as replaceable plugins—but practical interoperability remains unproven.
xAI says Grok 4.6 can sustain ambitious work across research, building, self-testing and revision—but the evidence remains supplier-reported.
Claude in Chrome can continue across desktop, web and mobile, but smoother handoffs do not remove the risk of hidden instructions inside web pages.
Spotify’s Xirp proposes a shared organisational-memory layer around coding agents, connecting services, owners, dependencies, decisions and session knowledge through Spotify Portal.
OpenAI says ChatGPT Work and Codex can import selected work from other agents and keep it updated, with important limits around compatibility, fidelity and availability.
Google DeepMind says SL2T turns on-device pose landmarks into server-translated text, bringing ASL input to Gboard and Live Transcribe on Pixel 11 first.
Koray Kavukcuoglu takes operational command of Google DeepMind while Demis Hassabis shifts towards Alphabet’s longer-term AGI and science direction.
Kimi pitches Slides as a research-to-presentation workspace where people can inspect and revise the generated structure, charts and drafts before downloading the deck.
OpenAI is bringing ChatGPT, ChatGPT Work and Codex into supported Linux desktop workflows, but the first preview has a tightly defined support matrix.
GLM-5 has arrived on Hugging Face with bilingual and mixture-of-experts tags, but independent evidence of its real performance is still missing.
Grok Bot pitches an early-beta shift from chat answers to delegated computer work, while reliability, security and results remain unproven.
ElevenLabs says Deutsche Telekom is moving AI voice assistance, translation, transcription and summaries into the telecom network.
MiniMax is pitching a connected creative workspace that divides a brief among specialist media agents and returns control to the creator at review checkpoints.
OpenAI says GPT-5.6-Cyber reached 95% on one internal task test and helped identify a V8 exploit chain, but wider results are mixed and access is restricted.
OpenRouter says Seedance 2.5 can combine image, video and audio references, generate up to 30 seconds, and extend or revise the result—with operational and privacy limits.
Anthropic’s Managed Agents now expose controls for session spend, inference geography, repository skills and stronger-model advice—with important limits still unstated.
Muse Glimmer brings a 30B agent model onto high-memory consumer hardware, but local does not mean lightweight, universally compatible or independently proven.
Grok’s new image model promises precise local edits, reusable visual worlds and more practical creative workflows—but its supplier-selected examples still need real-world testing.
Anthropic says its rewritten biology safeguard lets more everyday health and learning questions reach Fable 5 while dual-use research still falls back to Opus 5.
Claude Code sessions can exchange bounded text updates while remaining separate in history, files, permissions and control.