AI & Tech Brief ⚡
AI stopped merely writing code this week and started designing biology's tools, rebuilding its own product surface, and testing how much filesystem trust builders will extend to an agent.
📌 Navigate
📊 Exec Summary
AI stopped merely writing code this week and started designing biology's tools, rebuilding its own product surface, and testing how much filesystem trust builders will extend to an agent.
Five things moved in AI/tech this week:
IGI's Doudna lab ships AI-designed genome editors that beat nature
an inverse-folding + evolutionary-constraint pipeline produced RNA-guided nucleases that outperformed wild-type TnpB while holding specificity, active across bacterial, plant, and human cells — the week's clearest AI × biotech signal.
OpenAI's Codex surges while Anthropic reverses course
GPT-5.6 Sol/Terra/Luna hit Bedrock GA topping the coding-agent leaderboard, Codex grew >10x to 7M users, and Anthropic made Claude Fable 5 permanent after an earlier API-only plan.
Moonshot ships Kimi K3
the first open 3T-class model (2.8T params), self-positioned as trailing only Fable 5 and GPT-5.6 Sol, resetting the self-hosted frontier cost calculus.
xAI's Grok Build CLI caught exfiltrating local files
it uploaded entire working directories including SSH keys and password databases; xAI disabled it, deleted retained data, and open-sourced the codebase in response.
Thinking Machines ships Inkling
Mira Murati's lab's first open-weights model (975B/41B active, 1M context) with day-one Transformers support, framed as a customization base, not a leaderboard leader.
The pattern: design over discovery, agents over chat, open weights over open transparency.
1️⃣ IGI's Doudna lab ships AI-designed genome editors that beat natural CRISPR enzymes
TL;DR: IGI researchers used an inverse-folding model paired with evolutionary-constraint modeling to generate RNA-guided nucleases that outperformed the wild-type TnpB enzyme while maintaining specificity, with activity confirmed across bacterial, plant, and human cells.
What happened
- The method combines an inverse-folding model (structure-to-sequence, developed by Chloe Hsu and Alex Rives, formerly Meta AI) with evolutionary-constraint modeling of the nucleic-acid interface — not a sequence-only language model.
- Roughly 1 in 4 of ~2,000 AI-generated proteins tested were confirmed as active nucleases in wet-lab experiments.
- Variants showed editing activity in all three cell types tested (bacterial, plant, human), and the team solved the first-ever 3D structure of an AI-designed functional RNA-guided nuclease.
- Work spans the Doudna, Jacobsen (UCLA), Cate, and Banfield labs, combining wet-lab and computational science, published in Science as "Structure and evolution-guided design of minimal RNA-guided nucleases" (DOI 10.1126/science.aed6123).
📊 Benchmarks (from IGI newsroom)
| Measure | Result | Context |
|---|---|---|
| Lab success rate for AI-generated TnpB variants | ~1 in 4 | of ~2,000 tested proteins confirmed as active nucleases |
| Cell-type coverage | 3 | activity shown in bacterial, plant, and human cells |
| Solved structures for AI-designed variants | 1 | first solved 3D structure for an AI-designed functional RNA-guided nuclease |
🔗 Primary source → AI-Assisted Technique Allows Scientists to Design New, Functional Genome Editors Beyond What Can Be Found in Nature
🔍 The non-obvious point
The mechanism is the story: this extends beyond natural sequence space by design, not by scaling a language model on more protein sequences.
- Petr Skopintsev framed it as programming custom properties — "we generated variants that outperformed the activity of the wild type TnpB nuclease, yet maintained the specificity" — positioning the pipeline as a repeatable design method, not a one-off hit.
- The newsroom stays silent on the hard parts: results are cell-based only (no in vivo/whole-organism outcomes), there is no off-target/specificity profiling beyond the headline trade-off claim, and no timeline to a therapeutic or agricultural candidate.
- GEN's independent coverage corroborates the same Science paper, framing the synthetic nucleases as widening the CRISPR toolkit for next-generation editing therapeutics — the calibration builders should hold is "designable space expanded," not "clinical-grade editors shipped."
👀 What to watch
- Whether the team (or others) generalize the pipeline to other nuclease systems and report in vivo editing with off-target data — the primary explicitly pitches it as a method others can adopt.
2️⃣ OpenAI's Codex surges past 7M users as Anthropic makes Claude Fable 5 permanent
TL;DR: GPT-5.6 Sol/Terra/Luna went generally available on Amazon Bedrock topping the coding-agent leaderboard, Codex adoption grew more than 10x in six months to 7M users, OpenAI is rebranding ChatGPT around Codex, and Anthropic made Claude Fable 5 permanent across Max and Team Premium after an earlier API-only plan.
What happened
- OpenAI shipped a three-tier release: Sol (flagship reasoning, with a max-reasoning-effort setting), Terra (balanced production), and Luna (fast/high-volume inference), now GA on Bedrock's next-generation inference engine at OpenAI first-party rates.
- AWS explicitly names coding agents, cybersecurity research, and genomics workflows as target Bedrock use cases for the release.
- Per Latent Space/AINews, Codex usage grew >10x in six months to 7M users, adding ~1M users in roughly a day, raising whether it is overtaking Claude Code in adoption.
- Simon Willison confirmed Anthropic made Claude Fable 5 permanent across Max and Team Premium plans at 50% of usage limits starting July 20, reversing an earlier API-only plan, reportedly in response to competition from GPT-5.6 Sol and Kimi K3; he also confirmed via binary inspection that Claude Code now runs on a Rust-rewritten Bun runtime with ~10% faster startup on Linux.
📊 Benchmarks (from AWS/OpenAI, GPT-5.6 Sol)
| Benchmark | GPT-5.6 Sol | Comparison |
|---|---|---|
| Artificial Analysis Coding Agent Index | 80 | 2.8 pts above next-best, using <half the output tokens/time at ~1/3 the cost |
| ExploitBench (cybersecurity research) | 73.5% | vs. 47.9% for GPT-5.5 at a comparable output-token budget |
| Agents' Last Exam (55-field long-running workflow) | 53.6 | 13.1 pts above next-best; leads by 11.4 pts at medium effort for ~1/4 the cost |
🔗 Primary source → OpenAI GPT-5.6 Sol, Terra, and Luna are now generally available on Amazon Bedrock
🔍 The non-obvious point
The competitive war is entirely secondary-sourced — the AWS post never names Anthropic or Claude Code, so the "Codex vs. Claude Code" framing comes from commentators, not the announcement.
- Ben Thompson argued OpenAI's move to rebrand Codex as the new ChatGPT interface signals a core product-identity shift away from general chat toward an agentic coding/work surface — the product, not the model, is being repositioned.
- Zvi Mowshowitz published a cross-source model-selection matrix (Sol+Codex vs. Fable+Claude Code) on intelligence, trustworthiness, and agentic-workflow fit — the practical read for builders picking a default agent this quarter.
- Anthropic's reversal is the tell: making Fable 5 permanent after planning API-only access, reportedly because of Sol and Kimi K3, shows pricing/access terms are now a competitive-response lever, not a static plan. Note AWS lists genomics workflows as a first-class Bedrock target — the agent war reaches directly into biotech-AI pipelines.
👀 What to watch
- The Claude Fable 5 permanent-access rollout at 50% of usage limits begins July 20 across Max and Team Premium — watch whether it holds or shifts again under continued Sol/K3 pressure.
3️⃣ Moonshot ships Kimi K3, the first open 3T-class model
TL;DR: Moonshot AI released Kimi K3, a 2.8T-parameter mixture-of-experts model it calls the world's first open 3T-class model, available via its apps and API now with full weights promised by July 27, 2026.
What happened
- K3 uses Kimi Delta Attention (KDA), Attention Residuals (AttnRes), and a Stable LatentMoE framework activating 16 of 896 experts, with MXFP4 weights / MXFP8 activations from quantization-aware training and a claimed ~2.5x scaling-efficiency gain over Kimi K2.
- It is live today via Kimi.com, Kimi Work, Kimi Code, and the Kimi API; full model weights are not yet released at publication.
- Moonshot self-reports case studies including a from-scratch GPU compiler (MiniTriton) and an autonomous 48-hour chip-design run, and recommends supernode configs with 64+ accelerators for inference efficiency.
- The company states K3 "still trails the most powerful proprietary models, Claude Fable 5 and GPT 5.6 Sol," while consistently outperforming other tested models.
📊 Benchmarks (from Kimi K3 tech blog)
| Measure | Kimi K3 | Context |
|---|---|---|
| BrowseComp (1M-token context, no context mgmt) | 90.4 | comparison figures for Fable 5, Opus 4.8, GPT-5.6 Sol, GPT-5.5 cited from vendor announcement pages |
| Parameters | 2.8T total (MoE, 16 of 896 experts) | first open model to reach 2.8T; ~2.5x scaling-efficiency gain over K2 |
| Kimi API pricing | $0.30 / $3.00 / $15.00 per MTok | cache-hit input / cache-miss input / output; >90% cache-hit rate on coding via Mooncake |
🔗 Primary source → Kimi K3 Tech Blog: Open Frontier Intelligence
🔍 The non-obvious point
The most striking move is the self-positioning: a Chinese lab openly benchmarks against — and cites the announcement pages of — Fable 5 and GPT-5.6 Sol, and ships a dedicated "Limitations" section flagging thinking-history sensitivity and "excessive proactiveness."
- Alberto Romero framed K3 as the first Chinese open-weight model to reach frontier-model quality, narrowing a compute/talent gap builders assumed only US labs could close.
- Simon Willison calibrated it as trailing Fable 5 and GPT-5.6 Sol but ahead of Opus 4.8 and GPT-5.5 on several evals — a genuine frontier-tier open model, not a leaderboard leader.
- Latent Space noted it lands priced competitively with Claude Sonnet 5, reshaping the compute/cost calculus for self-hosted frontier-class inference — but the primary carries no independent third-party verification and no training-data or licensing disclosure.
👀 What to watch
Full open weights are promised by July 27, 2026
the release and first independent benchmarks are the real test of the self-reported standing.
4️⃣ xAI's Grok Build CLI caught uploading SSH keys, then open-sourced in response
TL;DR: xAI's Grok Build (grok CLI) was found uploading users' entire working directories — including SSH keys and password-manager databases — to xAI cloud storage; xAI disabled the behavior, deleted retained data, and open-sourced the codebase under Apache 2.0 in response.
What happened
- Two independent sources (Simon Willison and The Pragmatic Engineer) confirmed the CLI uploaded entire local directories including SSH keys and password databases before xAI disabled the behavior following user backlash.
- The public repository is xAI's remediation artifact: the Grok Build CLI/TUI Rust source, synced from the SpaceXAI monorepo, made public under Apache 2.0.
- The repo ships prebuilt binaries for macOS, Linux, and Windows, documents build-from-source requirements (Rust toolchain, DotSlash), and supports the Agent Client Protocol (ACP) for editor embedding.
- A SECURITY.md file exists in the repository, but no public postmortem describes the data-exposure scope, how much data was retained before deletion, or which users were affected.
📎 Technical identifiers (from the Grok Build repository)
Grok Build (grok CLI)
Rust source, periodically synced from the SpaceXAI monorepo
Apache 2.0 license
the open-sourcing is the disclosed response to the data-handling issue
Agent Client Protocol (ACP)
for embedding the agent inside editors
🔗 Primary source → GitHub — xai-org/grok-build: SpaceXAI's coding agent harness and TUI
🔍 The non-obvious point
Open-sourcing as apology is the move here — the repo is the remediation, not a feature launch, and it arrives without a postmortem.
- The concrete vendor-risk lesson: agentic coding CLIs run with full filesystem access, and a default telemetry/upload path swept up SSH keys and password databases — the highest-value secrets on a developer machine.
- The disclosure gaps matter operationally: no scope, no affected-user list, no changelog entry tied to the disabled upload behavior means teams can't independently assess exposure.
- The Pragmatic Engineer independently corroborated the same incident, which is what elevates it from a single-source rumor to a confirmed vendor-trust signal for anyone evaluating AI coding agents.
👀 What to watch
- Whether xAI publishes a postmortem quantifying data retained and users affected — absent that, treat the open-sourcing as partial remediation, not closure.
5️⃣ Thinking Machines ships Inkling, its first open-weights model
TL;DR: Mira Murati's Thinking Machines Lab released Inkling, its first open-weights model — a 975B-parameter (41B active) multimodal MoE with a 1M-token context — under Apache 2.0 with day-one Hugging Face Transformers support, framed as a customization base rather than a frontier leader.
What happened
- Inkling: 975B total / 41B active parameters, 1M-token context, pretrained on 45 trillion tokens of text/image/audio/video; a smaller Inkling-Small (276B total / 12B active) shipped as a preview.
- Both are Apache 2.0, with weights on Hugging Face (original + an NVFP4 checkpoint for NVIDIA Blackwell) and availability on Tinker for fine-tuning at 64K/256K context (50% discount for a limited time).
- The model was trained with controllable "thinking effort" — chain-of-thought grew more compressed and telegraphic over 30M+ RL rollouts without hurting final-answer quality.
- Launch ecosystem partners include TogetherAI, Fireworks, Modal, Databricks, and Baseten for API access, plus RadixArk, Inferact, Lightseek, Unsloth, and Hugging Face for inference; the company states Inkling "is not the strongest overall model available today, open or closed."
📊 Benchmarks (from Thinking Machines newsroom)
| Benchmark | Inkling | Inkling-Small |
|---|---|---|
| AIME 2026 (effort=0.99) | 97.1% | 95.1% |
| GPQA Diamond | 87.2% | 88.3% |
| SWEBench Verified | 77.6% | 77.4% |
| Terminal Bench 2.1 (best harness) | 63.8% | 52.7% |
🔗 Primary source → Inkling: Our open-weights model
🔍 The non-obvious point
Thinking Machines is explicitly not competing on raw capability — it positions Inkling as a customization-friendly base (multimodal, efficient thinking, Tinker fine-tuning), conceding it isn't the strongest model, open or closed.
- Note Inkling-Small edges the larger model on GPQA Diamond (88.3% vs. 87.2%) — a signal that the small preview is the more interesting artifact for cost-constrained builders.
- Simon Willison flagged that the model card and training-data documentation disclose very little about training-data sourcing — a governance/provenance signal for anyone weighing open-model licensing risk.
- The launch makes no comparison to Kimi K3, the other major open-weights release of the same week — leaving builders to run the head-to-head themselves; Hugging Face shipped day-one Transformers support (v5.14.0), so it drops straight into existing pipelines.
👀 What to watch
- Whether independent evals and provenance scrutiny confirm the customization pitch — with weights live on Hugging Face, third-party benchmarks and training-data questions land fast.
📊 The pattern
The frontier moved in four directions at once this week: AI designed the tools of biology rather than just describing them, coding agents became the product surface (Codex as the new ChatGPT, Fable 5 access as a competitive lever), open weights scaled to 3T-class while training-data transparency lagged behind, and a data-exfiltration incident showed vendor trust is the real gate on agent adoption. Capability is no longer the scarce input — provenance, trust, and product identity are. Design over discovery, agents over chat, open weights over open transparency.
👀 Watchlist
Kimi K3 full weights (July 27, 2026)
the promised open-weights drop and first independent benchmarks will test Moonshot's self-reported "trails only Fable 5 and Sol" claim.
Claude Fable 5 permanent access (from July 20)
rollout at 50% of usage limits across Max and Team Premium; watch whether terms hold under continued GPT-5.6 Sol and Kimi K3 pressure.
Grok Build postmortem
xAI has open-sourced the CLI but published no incident writeup quantifying data retained or users affected; its absence is itself a vendor-risk signal.
IGI pipeline generalization
the paper pitches the design method as portable to other nuclease systems; watch for in vivo editing results and off-target profiling.
Inkling provenance scrutiny
with Apache 2.0 weights live on Hugging Face, independent evals and training-data questions will surface quickly.
📎 Sources
Sources of truth
Click to verify or go deeper.
| Source | Title | URL | Date |
|---|---|---|---|
| Innovative Genomics Institute | AI-Assisted Technique Allows Scientists to Design New, Functional Genome Editors Beyond What Can Be Found in Nature | https://innovativegenomics.org/news/ai-designed-functional-genome-editors/ | 2026-07 |
| AWS Machine Learning Blog | OpenAI GPT-5.6 Sol, Terra, and Luna are now generally available on Amazon Bedrock | https://aws.amazon.com/blogs/machine-learning/openai-gpt-5-6-sol-terra-and-luna-are-now-generally-available-on-amazon-bedrock/ | 2026-07 |
| Moonshot AI (Kimi) | Kimi K3 Tech Blog: Open Frontier Intelligence | https://www.kimi.com/blog/kimi-k3 | 2026-07 |
| xAI | GitHub — xai-org/grok-build: SpaceXAI's coding agent harness and TUI | https://github.com/xai-org/grok-build | 2026-07 |
| Thinking Machines Lab | Inkling: Our open-weights model | https://thinkingmachines.ai/news/introducing-inkling/ | 2026-07 |
| Hugging Face | Transformers v5.14.0 release notes (day-one Inkling support) | https://github.com/huggingface/transformers/releases/tag/v5.14.0 | 2026-07 |
Commentary we read
| Author / outlet | Title | URL | Date |
|---|---|---|---|
| GEN (Genetic Engineering & Biotechnology News) | AI-Designed Synthetic CRISPR-Like Nucleases Show Activity in Cells | https://www.genengnews.com/topics/genome-editing/ai-designed-synthetic-crispr-like-nucleases-show-activity-in-cells/ | 2026-07 |
| Zvi Mowshowitz (Don't Worry About the Vase) | Better Call Sol: The Workhorse | https://thezvi.substack.com/p/better-call-sol-the-workhorse | 2026-07 |
| Latent Space / AINews | Codex usage up 10x in 6 months | https://www.latent.space/p/ainews-codex-usage-up-10x-in-6-months | 2026-07 |
| Ben Thompson (Stratechery) | The OpenAI Super-App: ChatGPT, Codex, Whither Chat | https://stratechery.com/2026/the-openai-super-app-chatgpt-codex-whither-chat/ | 2026-07 |
| Simon Willison | Anthropic makes Claude Fable 5 permanent | https://simonwillison.net/2026/Jul/18/claude-make-fable-5-permanent/#atom-everything | 2026-07-18 |
| Simon Willison | Kimi K3 | https://simonwillison.net/2026/Jul/16/kimi-k3/#atom-everything | 2026-07-16 |
| Latent Space / AINews | Kimi K3 2.8T-A50B: the largest open model | https://www.latent.space/p/ainews-kimi-k3-28t-a50b-the-largest | 2026-07 |
| Alberto Romero (The Algorithmic Bridge) | Moonshot Is Chinese But Its AI Models | https://www.thealgorithmicbridge.com/p/moonshot-is-chinese-but-its-ai-models | 2026-07 |
| Simon Willison | Grok Build | https://simonwillison.net/2026/Jul/15/grok-build/#atom-everything | 2026-07-15 |
| The Pragmatic Engineer (Gergely Orosz) | The Pulse: Grok's CLI caught uploading local files | https://newsletter.pragmaticengineer.com/p/the-pulse-groks-cli-caught-uploading | 2026-07 |
| Simon Willison | Inkling | https://simonwillison.net/2026/Jul/16/inkling/#atom-everything | 2026-07-16 |
| Latent Space / AINews | Thinky's Inkling 975B-A41B | https://www.latent.space/p/ainews-thinkys-inkling-975b-a41b | 2026-07 |