84 - Google Puts Agents on the Rust Rewrite.

Posted on

The Rust rewrite has stopped being a debate and become an engineering programme with numbers attached. Google says its agents are porting C and C++ to Rust, up to an 800,000 line kernel. A team three years into its own migration shares the dashboards, including the Tokio mistake that cost it a 3.7 GiB pod. And the compiler you build all of it with just got 4.57% faster in two months.

Featured sponsor
SvixSend webhooks without building the infrastructure. Svix is the enterprise-ready webhooks service, and its core is built in Rust. Retries, signature verification, fan-out, rate limiting, and observability arrive as an API you can drop in within minutes, so your team ships product instead of maintaining webhook plumbing. Start sending at svix.com →

Google Puts Agents on the Rust Rewrite

The most interesting paragraph in Google's Gemini 4 Argon announcement is not about benchmarks. It is about Rust. Google says Argon agents are migrating C and C++ codebases to Rust across the company, from tens of thousands of lines in core libraries like re2 and libgav1 up to more than 800,000 lines for the Fuchsia Zircon kernel.

The libgav1 example is the concrete one. Starting from an existing Rust port of Google's AV1 video decoder, agents replaced 32,000 lines of SIMD code by running rounds of profile-guided experiments and studying the compiler's output, until they had safe Rust that the compiler vectorises on its own. The result decodes 2.7 times faster than the earlier Rust port, with identical video output, closer to the hand-optimised C++.

Google is careful about the rest. Given how critical these systems are, the rewrites go through automated and manual auditing, emulation testing and review before production, and Argon itself is only rolling out to a set of trusted cyber defenders for now. The r/rust discussion was split between excitement and developers who have reviewed agent-written code and found it works but is hard to maintain.

Takeaways:

  • A company of Google's size now treats C and C++ to Rust as a migration it can scale, not a one-off rewrite
  • Safe Rust that auto-vectorises beat a hand-written SIMD port by 2.7 times, which says as much about the compiler as about the agents
  • Maintainability of agent-written code is the open question, and the community is right to keep asking it

Three Years of Rust Migration, and the Tokio Mistake

Most migration stories end at the cutover. Stephen Blum's three-year write-up shows fifteen dashboard panels, and every one of those dashboards existed before the migration started, so the before and after are real measurements. His team moved latency-sensitive services from Python, Go, JVM and C to Rust, one service and one region at a time, running old and new side by side.

The headline numbers: replacing an nginx-based load balancer with Pingora, Cloudflare's Rust proxy framework, took average balancer time from 600ms to 101ms on the same hardware and traffic. Publish latency on the hottest path went from about 350µs to about 50µs, and the variance dropped with it, which mattered more for their latency guarantees than the average.

Two of the panels compare Rust with Rust. A redesign of the Presence service cut peak pod memory about six times, a reminder that the first Rust release is a baseline, not the finish line.

The best section is the failure. They spawned a task per request with no cap, and during bursts Tokio queued tasks that each held buffers while waiting on I/O. A 100 MiB pod climbed to 3.7 GiB. Leak detection found nothing, because the memory was live. Scaling out only gave the backlog more places to grow. The fix was backpressure at the ingest edge: a cap on in-flight work and a bounded queue.

Takeaways:

  • Instrument before you migrate, or you cannot prove the result
  • Async tasks are cheap in CPU, not in memory: put a limit somewhere in the path
  • Treat the first Rust version as a baseline, iterate from there

rustc Got 4.57% Faster in Two Months

Nicholas Nethercote's latest compiler performance update covers 29 July to 28 September, and the summary is unusually strong. The mean wall-time reduction across the benchmark suite was 4.57%, and 555 of 629 benchmark measurements improved.

The wins came from many people. Nikita Popov's upgrade to LLVM 23 alone gave a mean 1.2% reduction. Jakub Beránek enabled PGO for Clippy, worth up to 18% on Clippy benchmarks. A new contributor, xmakro, landed a string of improvements, including one that cut mean cycle counts by 1.58%. At one point so many performance PRs were waiting to merge that they were bundled into a rollup of ten.

Nethercote's own favourite is a change to the traversal order of the compiler's dataflow analyses. For most code it changes nothing, but cranelift-codegen has one function with over 18,000 basic blocks. A borrow checker analysis that needed 1.5 million block visits to reach a fixpoint now needs 90,000, a 30% wall-time reduction for a check build of that crate.

There is a cost on the horizon too. The new borrow checker, Polonius Alpha, and the new trait solver are both enabled on Nightly, and both are slower for a minority of crates, including serde. Work on that is ongoing, and Nethercote is starting at Hexcat to work on the compiler performance optimisations project goal.

Takeaways:

  • 4.57% mean wall-time reduction in two months, with 555 of 629 measurements improving
  • LLVM 23 and PGO for Clippy were the big single changes
  • Polonius Alpha and the new trait solver are on Nightly, with known regressions being worked on
  • Nethercote moves to Hexcat to work on the compiler performance project goal

PolyXOR128, a Fast Hash With a Machine-Checked Proof

Fast hashes and provable hashes are usually different hashes. Orson Peters, the author of pdqsort, has released PolyXOR128, a 128-bit universal hash that tries to be both. With a key that is random and independent of the input, the chance that two inputs of up to n bytes collide is at most (n/4096 + 3) / 2^128, and that proof is formalised in Lean.

On speed, he reports up to 130 GB/s on a Ryzen 9950X and 60 GB/s on an Apple M2 Pro, which makes it, to his knowledge, the fastest 128-bit hash, ahead of most 64-bit and non-secure hashes too. The trade-off is hardware: it needs accelerated carryless multiplication, which every modern desktop and server CPU has, and the portable fallback is far too slow to use.

It is a universal hash family, in the same category as Poly1305 and GHASH, not an unkeyed hash like SHA-256. The key has to be secret, random and independent of the message, and Peters is explicit about that in the r/rust thread.

The Rust angle is his reason for not writing it in C. Cargo made it easy to depend on a hardware-accelerated AES implementation, and target_feature plus is_x86_feature_detected made runtime dispatch much simpler than in C, with faster code as a result.

Takeaways:

  • A 128-bit hash for file and stream integrity, with a collision bound proven in Lean
  • Up to 130 GB/s, if your CPU has carryless multiplication
  • It is keyed: a secret, random, independent key is part of the security claim
  • Cargo and target_feature are why it is a Rust crate and not a C library

Snippets


We are thrilled to have you as part of our growing community of Rust enthusiasts! If you found value in this newsletter, don't keep it to yourself — share it with your network and let's grow the Rust community together.

👉 Take Action Now:

  • Share: Forward this email to share this newsletter with your colleagues and friends.

  • Engage: Have thoughts or questions? Reply to this email.

  • Subscribe: Not a subscriber yet? Click here to never miss an update from Rust Trends.

Cheers,
Bob Peters

Want to sponsor Rust Trends? We reach thousands of Rust developers biweekly. Get in touch!