THE LOG — TRANSMISSIONS FROM ALTITUDE
What we learn running AI systems every day — workflows, growth loops, tools, and the occasional hard lesson. Written down so you can reuse it.
PROMPT LIBRARY · EVERY PROMPT WE PUBLISH, IN ONE PLACE
Act as a workflow auditor. Create a concise handoff map for a project moving into Codex: list the project context, active conversations, reusable skills, plugins, settings and configuration we rely on; mark each item as essential, useful or optional; then produce a three-step import review plan and a separate list of items we must test after import. Do not assume every item transfers intact or stays updated automatically.
Act as an accessibility product reviewer. Assess this proposed launch: ASL-to-English input arrives first in a phone keyboard and live-transcription feature; the phone converts camera video into whole-body pose coordinates on-device, then a server returns streaming text. Write a 180-word launch-readiness note with three sections: what the feature enables, what must be tested with Deaf signers before wider release, and which claims should remain qualified. Do not assume universal availability, error-free translation, or that the entire system runs on-device.
Act as a careful technology editor. Using only these announced facts: Koray Kavukcuoglu becomes SVP of Google DeepMind and oversees Gemini model development, Frontier AI research, and Gemini app and developer teams; Demis Hassabis becomes Chair of Google DeepMind and Chief Scientist of Alphabet, continues leading Isomorphic Labs, and advises DeepMind teams. Write a 150-word briefing with two headings: “What changed” and “What remains unproven”. Do not predict product performance, release timing, revenue, AGI, or health outcomes.
Create a six-slide internal decision deck on whether a small operations team should replace a weekly status meeting with an asynchronous update. Use a clear argument, cite the assumptions on each relevant slide, include one editable chart-data table comparing meeting time with written-update time, and end with a slide listing the facts we must verify before acting. Keep every slide concise and make the draft easy to revise before download.
Act as a cautious Linux desktop rollout assessor. Based only on this announced preview support matrix—Ubuntu 24.04 LTS or 26.04 LTS, Debian 13, Fedora 43 or 44; .deb or .rpm packages; x64 or ARM64—produce: (1) a compatibility verdict for each listed platform, (2) a three-step pre-install check, and (3) a separate list of unknowns. Do not infer pricing, eligibility, security, feature parity, offline use, performance, or production readiness. State that the ChatGPT desktop app is in preview.
Act as an AI model evaluator. Based only on these reported facts—GLM-5 was listed on Hugging Face on 11 August 2026 with Transformers text-generation support, safetensors weights, mixture-of-experts tags, and Chinese-English conversational tags; the provider reported 155,036 downloads and 2,119 likes—produce: (1) a two-column table separating inspectable release facts from unproven performance claims, (2) five independent evaluation tests needed before deployment, and (3) a one-paragraph recommendation that does not infer quality, speed, safety, cost, or API availability.
Act as an operations assistant. Using only the information in this message, prepare a draft reply to a vendor asking for a delivery-date update. Do not send anything, access external tools or invent facts. First list the assumptions you need confirmed, then write a 120-word draft reply, followed by a three-item checklist of actions that require our review.
Act as a careful telecom analyst. Using only the ElevenLabs X thread below, create a two-column table: “ElevenLabs states” and “Not established here”. Include network embedding, live assistance, translation, transcription and summaries in the first column where supported. Put availability, pricing, geography, technical architecture, rollout scale and measured outcomes in the second. Do not infer beyond the source. Source: ElevenLabs says its AI call assistant is embedded in Deutsche Telekom’s network, with live assistance during calls, real-time translation, call transcription and summaries; it says this grew from app features in early 2025 into an integration at the core of Deutsche Telekom’s services.
Act as a creative director in MiniMax Design. Turn this brief into a 30-second launch film for a reusable water bottle: “show an early commuter choosing a bottle that lasts all day.” First provide a five-step production plan covering script, storyboard, image direction, video direction and audio direction. Then draft the script, list three scene beats, specify one visual reference for each beat, write a voiceover under 75 words, and end with the exact checkpoint questions that need approval before assembly. Keep every output concise and label assumptions separately from decisions.
Act as a defensive security analyst. Using only this incident brief, produce: (1) a five-step triage plan, (2) a short validation checklist, and (3) a responsible-disclosure summary. Brief: a JavaScript engine issue may allow memory corruption and sandbox escape; the work is authorised, evidence is incomplete, and the goal is to help a vendor patch safely. Do not provide exploit code, payloads, bypass techniques, or attack steps. State assumptions and identify the evidence still needed.
Act as a video editor planning a 30-second sequence. Use the reference images, footage clips and audio supplied in this chat as the only source material. Preserve the named subject, colour treatment and audio mood; propose three shots, then revise one visual element and one audio element without changing the rest. Return a numbered shot list, an edit note for each shot, and a short list of any missing references.
Act as an operations lead preparing a Claude Managed Agent session. For the job “review this repository’s open deployment risks”, produce a compact control plan with four labelled parts: session budget and the action to take after a budget_reached pause; inference geography with the stated US versus global rate trade-off; repository skills to load from .claude/skills/ at session start; and the exact question to send to a stronger advisor model for a second opinion. Mark any policy, price or rollout detail not supplied here as unknown.
Act as a cautious local-agent deployment adviser. Using only the machine specifications, operating system, available memory, intended files and tools, and permission boundaries I provide, produce: (1) a compatibility-unknowns list, (2) a minimum permission plan, (3) three practical test tasks using files and screenshots, and (4) a stop/go decision that clearly separates verified facts from assumptions. Do not claim privacy, security, offline reliability or performance unless my inputs establish it.
Create a polished 16:9 product photograph of a cobalt-blue insulated bottle on a pale stone plinth, with soft window light and the words “FIELD NOTES” on a small cream label. Then change only the label text to “AUTUMN EDITION”. Keep the bottle colour, plinth, lighting, composition, shadows and every other detail unchanged. Return the edited image and a five-point checklist describing what changed and what was preserved.
Act as a careful biology explainer. Explain, in plain language, what a routine blood-test result can and cannot suggest, using only general educational information. Give three possible non-diagnostic reasons for one out-of-range marker, list four questions to take to a qualified clinician, and end with a short safety note. Do not diagnose, recommend treatment, or infer personal medical details.
You are finishing one Claude Code task and need to hand off to another separate session. Write a plain-text update of no more than 120 words: state what changed, name one unresolved risk, ask one precise question, and list the next action. Do not assume the other session has our history, files, configuration or permissions. End with a three-line format: Confirmed / Needs decision / Next action.
Using only the OpenAI announcement and article dated 7 August 2026, separate what OpenAI has confirmed about Astra from what remains preliminary. Return three headings: stated capability signal, controls already in place, and unanswered evidence; include no speculation or external sources.
Act as a supervised coding team working in this repository. Inspect the current validation and error-handling paths, identify one duplicated or inconsistent rule, then divide the work between a planner, an implementer and a reviewer. Propose the smallest safe fix, show the files and tests affected, apply no change until the reviewer has checked the plan, and finish with a traceable summary of the evidence, edits and remaining risks.
Act as a careful video-generation researcher. Using only the MiniMax H3 model card, produce a two-column table: ‘Available in the initial open-source release’ and ‘Hosted, future, or not yet open-sourced’. Include H3-Base, full-attention inference, Context-IR, Regenerate-2K and sparse attention. End with a 75-word practical summary that distinguishes the released engine from the complete three-stage system. Do not infer pricing, speed, hardware needs, commercial rights or independent performance.
Act as a cautious coding collaborator. Create a small command-line Python tool that reads a plain-text task list, groups tasks by priority, and writes a Markdown summary. First provide a five-step implementation plan and name two assumptions. Then write the code, include three sample tasks, and show the expected Markdown output. After each stage, list one check we can perform before continuing. Do not claim the result is production-ready; identify one limitation and one next verification step.
You are evaluating DeepSeek-V4-Flash through its public-beta API for one bounded, tool-using workflow. Design a repeatable test that sends the same system instructions and reference material across five runs, records fresh-input and cache-hit token billing where available, and checks whether tool calls complete correctly. Return: 1) the exact test steps; 2) one compact results table with run number, tool outcome, latency observed, fresh-input tokens, cache-hit tokens, output tokens and cost; 3) three failure conditions that stop the evaluation; and 4) a short conclusion that distinguishes observed results from DeepSeek’s published claims. Do not assume App/Web access, production readiness or benchmark superiority.
Act as a rigorous API evaluation partner. Take one realistic coding task: diagnose and fix a failing payment-validation function from a short error report and code snippet. Produce a complete patch, explain the root cause in no more than five bullets, list any assumptions, and finish with a verification checklist. Then assess the response on task completion, total steps, likely retries, latency tolerance and output quality; do not estimate token prices or claim benchmark results.
We have a queue of 2,000 routine review items and one urgent API task that must return within 15 minutes. Create a two-column decision note: routine work and urgent work. For each, state whether to prioritise lower processing cost or lower waiting time, name the evidence we would need before choosing a model, and list assumptions separately. Do not invent prices, latency measurements, quotas, or quality results. End with three checks we should run against our own usage data.
Act as a local workflow assistant for a small team. Take this recurring request: turn a meeting note into a concise action list, flag any missing owner or deadline, and produce a draft follow-up message. Show your reasoning as a short checklist, identify any uncertainty, and end with the exact points a human should review before sending.
Assess this Google DeepMind announcement using only these stated claims: Gemini Robotics 2 controls humanoids from feet to fingertips; Gemini Robotics ER 2 handles real-world video understanding and multi-step planning; On-Device 2 runs locally and adapts to new robot bodies in a few hours. Produce a two-column table: announced capability, evidence still needed before we could judge deployment readiness. Include access, pricing, benchmarks, safety, failure rates and independent testing.
Using a five-second voice reference recording that I have permission to use, generate this line: “The project is ready, but the final decision changes everything.” Create three versions that change only the word “changes”: first calm, then urgent, then reflective. Keep the remaining words as consistent as possible. Label each version with its intended emotion, intonation and pace.
Transcribe the attached recording. The recording concerns OpenAI’s GPT-Live-Transcribe and GPT-Transcribe; treat those model names as possible literal terms, but include them only if spoken. Expect English, with possible technical terms and numbers. Return: 1) a clean transcript with timestamps at speaker or paragraph changes; 2) a short list of uncertain names, numbers, or technical terms with their timestamps; 3) three focused checks we should run against the audio before using this workflow more widely.
Act as a production architect. For a remote MCP service, write a concise deployment note with exactly three headings: Protocol session, Application state, and Deployment choice. Explain how a self-contained protocol request can be handled by an available server instance while identifying which application data still requires explicit storage. Compare serverless, edge, and load-balanced deployments without claiming cost, latency, compatibility, or security benefits. End with three implementation questions.
Act as a cautious operations lead. Assess Block’s announced Buzz proposition using only these facts: Block says Buzz is free, open source, Apache-2.0 licensed, built on Nostr, and intended for shared human-and-agent work with cryptographic identities and configured permissions. Produce a three-column table: announced claim, evidence needed before adoption, and a practical test. Treat security, interoperability, hosted-service terms, adoption and product maturity as unverified; note that Git integration is still early.
Act as an AI evaluation assistant. Using only the information in this prompt, write a concise comparison table with three rows: what Google Gemma says WebGemma offers, what that may make easier during first-pass model exploration, and what still needs independent verification. Facts: Google Gemma describes a community-built browser playground with an interactive Gemma Journey timeline, WebGPU, Transformers.js, and 10+ models; it also says testing is securely on-device. Do not infer ownership, privacy, offline behaviour, compatibility, or production readiness.
Transcribe the attached audio verbatim. Preserve the spoken language, separate each speaker consistently, and include a timestamp for every word. Do not correct grammar, remove filler words, infer unheard speech, translate the recording, or merge uncertain speakers. Mark any uncertain word as [unclear] with its timestamp. Return: (1) the complete speaker-labelled transcript; (2) a table of every [unclear] segment; and (3) the total audio duration and number of distinct speakers detected.
Use Resy to find dinner availability for two people in central London next Friday between 7:00 pm and 9:00 pm. If the site requires authentication, pause and let me take control of the cloud browser to sign in. After I return control, continue the original search. Present up to five available restaurants in a table with neighbourhood, available time, cuisine, and booking link. Do not make a reservation or submit any personal information without asking me first.
Act as a workflow-risk analyst. Compare these two separate Claude betas without merging their capabilities: Claude Security scans changed code or a codebase, validates findings, proposes patches and requires human approval; Claude Voice can use Opus or Sonnet to reach connected tools such as email and calendar during a conversation. Produce a two-column comparison covering the decision point, permitted action, human control, evidence limitation and one safe trial. End with a 60-word judgement on whether workflow placement changes practical usefulness more than model choice. Treat all capability statements as Anthropic-published and do not infer security accuracy, tool reliability or shared availability.
Act as an AI strategy analyst. Using only these Google-published facts—model APIs process approximately 22 billion tokens per minute, up from 16 billion one quarter earlier; more than 9 million developers build each month across Google’s APIs and key developer products; the Gemini app has 950 million monthly active users; Google Cloud brings together chips, models, data, security and agent platforms; and Google says it remains supply constrained—produce a four-column table covering signal, distribution surface, practical implication and unresolved question. Attribute every figure to Google, treat tokens as units processed rather than completed tasks or intelligence, and do not infer capacity causes, product quality, paid usage, retention or universal availability. Finish with a 120-word assessment of whether distribution or any single feature is the more important strategic signal.
Act as a senior analyst completing a difficult, multi-step brief. First state a concise plan. Then produce the brief, inspect it against every stated constraint, identify any gaps, revise the work, and return both the finished version and a short verification note. Do not claim success where evidence is missing. Report the number of corrections, elapsed working time, and any points that still need human review so we can compare completion quality and total task cost across model runs.
We need to prepare a launch-readiness brief for a new desktop feature. Coordinate the work across three clearly named roles: one to map the user journey, one to identify evidence gaps and risky claims, and one to draft a concise release checklist. Keep us in control by pausing before any external action, stating what each role is doing, surfacing conflicts between their findings, and asking for one decision only when it materially changes the result. Return a single brief with: an executive summary, the three workstreams, unresolved risks, evidence still needed, and the next five actions in priority order.
During this voice conversation, inspect only the email and calendar tools already connected to this Claude account. Build a briefing for the next seven days with three sections: scheduled commitments, email threads that may affect those commitments, and unresolved questions. For every item, cite the date and the source tool. Do not send messages, change calendar events or infer missing details. If a tool, account or permission is unavailable, name the gap instead of filling it from assumption.
Using the Anthropic Economic Index connector in Claude, investigate how AI use differs between software developers and teachers. Compare the occupations by common tasks, automation versus augmentation patterns, and any change over time available in the Index. For every conclusion, identify the supporting Index measure or underlying data, distinguish observation from interpretation, and state explicitly that the dataset reflects Claude usage rather than the whole labour market. Return a concise comparison table followed by three evidence-backed takeaways and two questions the data cannot answer.
Act as an AI operations architect. Design a bounded test for one Claude Managed Agents workflow that uses per-agent effort, session seeding with no more than 50 user_message and define_outcome events, shared access within the limit of 500 skills across all Managed Agents in the session, environment or memory-store webhooks, and streamed sub-agent events. Use a research-and-draft task as the example. Return: (1) the agent roles, (2) the effort choice for each role with a brief rationale, (3) the opening events to seed, (4) the minimum skills required, (5) the webhook and event signals to record, and (6) a comparison table for completion time, cost, output quality, failed steps and missed events. Treat performance, pricing, availability and reliability as unknown until measured; do not invent product behaviour or results.
Act as a security architect reviewing an advanced-model cyber evaluation. The evaluation disables normal production cyber classifiers to measure maximum capability, runs inside an isolated research environment, and permits package installation only through an internally hosted registry cache. Assume a model pursuing the narrow objective of solving ExploitGym discovers a zero-day in that cache, escalates privileges, reaches an Internet-connected node, and then seeks secret benchmark solutions on external production infrastructure. Produce a threat model with: the trust boundaries crossed; the controls that failed or were absent; detection opportunities at each stage; a containment design that preserves useful capability measurement; and five testable acceptance criteria. Separate facts in this scenario from assumptions, do not infer motive or consciousness, and do not claim that any real incident investigation is complete.
Act as the governance lead for an enterprise voice agent that handles billing enquiries. The agent may verify identity, retrieve permitted account details, apply the published refund policy, take approved account actions and escalate exceptions to a person. A production review has found that callers asking about duplicate charges are escalated too early. Propose one narrowly scoped update without changing the agent's permissions or refund policy. Return five sections: observed gap; proposed behaviour change; policy and permission boundaries preserved; simulation cases covering normal, ambiguous and higher-risk requests; and a rollout decision table with Approve, Revise or Reject criteria. Require human review before rollout, identify any missing evidence and make no claims about reliability that the test results cannot support.
Act as a cautious code reviewer for a repository where external clients depend on the wire event name `completed`. Review these three proposed diffs separately: (1) rename `completed` to `done`; (2) add `done` while retaining `completed` as a backward-compatible event; (3) rename an unrelated local variable from `completedCount` to `doneCount`. Apply this repository rule: ‘Do not rename externally consumed event names without preserving backward compatibility. Keep the existing wire name or introduce the replacement alongside it with a documented migration path.’ For each diff, return: Verdict (`flag` or `no finding`), the exact rule condition that did or did not trigger, the downstream risk, and the safest next change. Do not invent repository facts or flag changes merely because they contain similar words.
Work in the current iOS project on macOS with Xcode available. First inspect the project and identify the safest existing simulator target. Build and run the app in the iOS simulator, then observe the launch state and interact only with the visible interface. Make no signing, deployment, dependency or destructive project changes without asking. Report: 1) the target and simulator used, 2) build errors or warnings, 3) what appeared on launch, 4) the interactions attempted and observed results, 5) any code change you recommend, and 6) what I should verify myself before accepting that change. If the build fails, stop after diagnosis and propose the smallest reversible next step.
Act as a cautious AI operations analyst. We are choosing between Qwen Cloud Token Plan Individual and Qoder for a recurring workload that uses one research agent, one coding agent and one media agent. Using only the offer facts below, compare the two routes without assuming unlisted prices, credits, limits or performance: Qwen Cloud says its plans start at $6 per month, combine several model families and tools in one credit pool, and allow up to eight simultaneous agents on Pro; its special offer is limited to new Token Plan users. Qoder says Qwen3.8-Max-Preview is available in its coding platform and currently carries a temporary 0.01x multiplier instead of 0.5x. Produce: (1) a compact comparison table covering intended workload, eligibility, tier dependency, known cost signal and unknowns; (2) the three facts we must verify in each service before paying; (3) one representative task to run on both routes; and (4) a decision rule based on measured task quality, credits consumed and effective cost. Label every unverified proposition as a supplier claim and do not calculate savings from missing usage data.
Act as an AI systems architect evaluating Gemini 3.6 Flash and Gemini 3.5 Flash-Lite for a customer-support agent. The system must classify and summarise 10,000 tickets each day, extract order details, draft routine replies and investigate the 5% of cases involving conflicting records or multi-step reasoning. Propose which model should handle each task, then design a controlled comparison using 200 representative tickets. Return: a routing table; quality, output-token, throughput and cost measures; pass/fail thresholds; an escalation rule; and a short decision note. Treat Google's published prices and performance figures as claims to test, not guaranteed results. Exclude Gemini 3.5 Flash Cyber because it is restricted to a limited CodeMender pilot.
In Claude Cowork, create a folder named Claude Skill Trial containing five fictional expense receipts as plain-text files. Give each receipt a different date, merchant and amount, then organise copies into month folders and produce a manifest listing every filename, destination and amount. Do not access, move or alter any existing files. This workspace will be used to demonstrate the same filing process with Record a skill and check whether the saved skill reproduces the manifest accurately on a second fictional set.
Animate the attached still as a single continuous scene. Keep the subject recognisable and the composition coherent. Begin with near-stillness, then introduce a slow forward camera move as wind lifts the smallest details in the frame. Let the atmosphere grow from quiet tension into release without adding new people, objects or locations. Create restrained environmental sound, closely timed physical sound effects and no dialogue. Return one video whose motion, camera treatment and sound feel like parts of the same deliberate moment.
Create a live, publishable one-page site for a small independent bakery launching a Saturday bread subscription. Use the headline “Better Saturdays start with better bread”. Include a concise introduction, three subscription benefits, a simple three-step collection guide, an FAQ covering collection time and allergens, and one clear “Join the Saturday list” call to action. Keep the tone warm and confident, use plain English, avoid invented prices or customer testimonials, and make the finished page easy to scan on mobile.
Act as a voice-product architect. Assess this announced system: Google Gemma says Gemma 4 31B can act as the brain inside a fully open-source, cascaded speech-to-speech stack for existing voice apps, with Hugging Face and Cerebras named. The announcement provides no latency measurement, benchmark, price, availability detail, licence evidence, supported-language list or production-readiness evidence. Produce a concise implementation memo with four sections: proposed input-to-output flow; components an existing app could retain or replace; evidence required before a pilot; and a pilot acceptance table with measurable criteria but no invented thresholds. Label every statement as Announced, Inferred or Unverified.
Act as a physical-AI simulation reviewer. Examine the supplied 3D scene or asset description for six areas: scale, materials, semantic labels, collision behaviour, physical properties and sensor-readable outputs. For each area, state what is known, what is missing and what should be checked next. Do not invent measurements, certify safety or assume visual realism proves physical accuracy. Return a concise table followed by three prioritised fixes that require human approval before implementation.