ZCode caught uploading entire repos to Alibaba Cloud, hackers chain a heap overflow into OpenAI's internal monorepo, and Bend wants AI to prove its code correct.
The trust surface of AI coding tools just shattered. Two stories this week — one about silent data exfiltration, one about chaining an image upload into internal repo access — are a wake-up call for anyone giving an AI agent write access to their codebase.
ZCode, the desktop coding agent built by Z.ai (the company behind GLM open-weight models, publicly listed on the Hong Kong Stock Exchange since January), was caught silently packaging entire workspaces — including complete .git histories — encrypting them, and uploading the archives to Alibaba Cloud.
Researcher ferstar reverse-engineered the client and found that a 345MB commercial workspace with 42,411 files produced a 313MB encrypted archive. The .git directory alone comprised 86.6% of the payload: LFS objects, commit history, logs. That means deleted API keys, unpushed branch names, internal hostnames — everything your git history has ever touched.
The encryption uses AES-256-CTR with an RSA-OAEP envelope. The private key lives exclusively on Z.ai's servers. You can't even inspect what was taken. The client maintains persistent connections to zcode.z.ai and two Aliyun OSS nodes, silently streaming your codebase to Alibaba Cloud's object storage.
The damning detail: UI toggles labeled "Optimize Experience" and "Repo Snapshot Indexing" did not stop the uploads. "Optimize Experience" only controls model training authorization — snapshot capture and upload continue regardless. "Repo Snapshot Indexing" only controls server-side indexing — local packaging and upload continue. The capture sidecar launches unconditionally at startup, requiring only a valid JWT. ferstar logged 62 capture events in a single session and 564 failed upload attempts. ferstar's summary thread hit 276,000 views within 13 hours. Z.ai's only known response came from an affiliated account: "hey I am sorry to let you find it."
This isn't a misconfigured telemetry endpoint. This is an envelope-encrypted full-repo exfiltration pipeline with no functional off switch. If you've run ZCode against a proprietary codebase, assume your complete repository lineage — every commit, every deleted secret, every unpushed experiment — is in Alibaba Cloud right now.
Bend dropped this week with a proposition that sounds absurd until you think about it for five minutes: a language where AI agents must produce formal proofs that their code satisfies declared invariants before it can merge. It hit 523 points on Hacker News and sparked 252 comments — roughly split between "this is the future" and "formal verification has been 10 years away for 40 years."
The pitch is four pillars: C speed, CUDA parallelism, Lean-class proofs, Python syntax. It compiles to native code, runs on CPU or GPU (they demoed 4,096 GPU cores), and includes a type checker that doubles as a proof checker — finishing in about a second versus the minutes Lean or Rocq typically need. Two academic papers underpin the system: BendTT (an affine dependent type theory) and BendRT (a parallel runtime for CPUs and GPUs).
The interesting part is LAWS.bend. You declare invariants as formal laws:
When an AI agent writes code that touches the game board, it must also produce a PROOF.bend that type-checks against those laws. If the proof fails, the merge is blocked. Their demo showed an AI adding a wrap-around feature to a game — without LAWS.bend the bug shipped; with it, the AI had to retry until it built a guard and proved the law holds. The Bend team's claim: "merging a bug is mathematically impossible — it is a theorem."
They frame it as "LAWS.bend is AGENTS.md backed by proof" — and that framing is sharp. We're already writing natural-language constraints for coding agents in AGENTS.md and CLAUDE.md files. Bend asks: what if those constraints were machine-checkable?
The practical limitations are real. Bend is early ("expect bugs"), works best on Linux and macOS backends, and writing proof obligations requires a skill set most teams don't have. Formal verification has historically failed to escape academia because the annotation burden exceeds the bug-finding benefit for human-authored code.
But the context has changed. When a human writes code, they can reason about it incrementally. When an AI agent writes code at scale, the verification burden falls entirely on review — and reviewers miss things. A proof obligation shifts the burden back to the generator. The AI doesn't need to be right on the first attempt; it needs to prove it's right before shipping.
I don't think Bend will replace Python or Rust. But the idea of machine-checkable invariants for AI-generated code is going to show up everywhere in the next two years. Watch for this pattern: agent generates code, tool checks formal properties, agent iterates until the proof holds. shadcn/lint (below) is already doing a lightweight version of this for design systems. The teams building coding agents should be watching this closely.
shadcn-ui/lint — An agent-first linter for Tailwind design systems. Define per-component rules (allow typography changes on CardTitle but block color overrides); when an AI agent breaks one, the error explains what's wrong and suggests a fix using your actual design tokens. Add the lint command to your AGENTS.md and agents run it after every change. Across 150+ evaluation runs, Sonnet 5, Haiku 4.5, Opus 5, and GPT 5.6 Terra all reached zero violations after one correction round (8/8 tasks completed each). Fixing with lint feedback cost 10–48% less than relying on rules alone. Supports ESLint 9.30+ and Oxlint 1.80+. 2.1K stars, MIT.
tigerless-labs/agent-memory — Long-term memory runtime for AI agents using plain Markdown as source of truth with local BM25 + optional vector retrieval fused via Reciprocal Rank Fusion. Claude Code and Codex CLI share one store, zero API keys required. The sleep-time Manage layer runs consolidation autonomously — distillation fires at conversation boundaries rather than relying on the agent's judgment. 52.9% accuracy on LongMemEval-S vs. 35.8% for MemCore (p=0.009) and 5.8% for no memory (p<0.001). Cross-host test across 9 writer/reader pairs improved net contribution from 2/36 to 13/36. 939 stars, MIT.
Bend — The proof-backed language from the Deep Dive above. Compiles to native, auto-parallelizes to GPU, and lets you declare formal laws that AI-generated code must satisfy before merge. Python-like syntax with a type checker that doubles as a proof checker (finishes in ~1 second vs. minutes for Lean/Rocq). Early-stage but conceptually the most important repo in this list.
Three threads converged this week into a single uncomfortable picture.
The OpenAI hack is technically fascinating — Claude Opus 5 producing a working ASLR bypass exploit within three hours of its public release, a heap overflow in an image decoder chained through Discourse, SSO, ChatGPT, Codex, and into an internal monorepo. The Hacktron researchers needed a genuine zero-day, an SSO misconfiguration, and serious exploitation skill.
But the ZCode story is the one that should keep you up at night, because ZCode just needed you to install it. No exploit required. Every AI coding tool you run has filesystem access to your entire project. Most have network access too. The attack surface isn't the vulnerability — it's the permission model we accepted without questioning.
Meanwhile, the unredacted NYT filings show that even the people building these systems know the current trajectory is unsustainable. Microsoft's own Director of Applied Science called it "the largest theft of labor in human history." OpenAI's head of ChatGPT wrote internally that their products pose an "existential threat" to publishers. When the builders are this candid in private, the industry's public posture of "this is just fair use" looks increasingly fragile.
Audit your coding tools' network traffic this weekend. You might not like what you find.
— Aaron, from the terminal. Watch what your tools watch.
Compare AWS Bedrock AgentCore, LangChain, and Alibaba AgentLoop for enterprise AI agents. Architecture, cost, and production trade-offs.
AI EngineeringTraditional SEO drives clicks through rankings; inference traffic converts through AI citations. Learn which strategy fits your content goals in 2026.
SEOModern SEO packages must cover AI citation optimization and video search. 56% of search activity happens through AI, while 82% of all internet traffic is video.
SEO