Sum over paths

View of a bridge from the Tadami Line, Japan
© 2024. Akshay Mishra

The capital goods of software engineering

Can an LLM write sudo?

A lathe. Source: Scammell, Henry B. (1897). Cyclopedia of Valuable Receipts: A Treasure-House of Useful Knowledge for the Every-Day Wants of Life (Public domain)

I find Arena AI’s “unboxing’ videos of newly launched AI models to be quite fun to watch. This morning I saw the one for GPT 6 Astra drop.

As they went through the examples one by one, the results seemed quite impressive, especially where they displayed an FPV open-world game with ostensibly “10 hours of gameplay” and also a D-day landing 3D scene. Those were built atop what I presume was three.js and some MCP-based Blender modelling. Being a game-developer, seeing this level of automation producing coherent results is fascinating indeed (and a bit unsettling!)

Now this is a pattern we have all become intimately familiar with over the past twelve-ish months : LLMs producing surprisingly good results through the orchestration of agentic workflows, making tool calls and using MCP hooks to drive software to desired results.

But it is worth taking a step back to acknowledge the elephant(s) in the room: the tools themselves.

Production = process + tools; LLMs (largely) learn process

There is a reason why LLMs thrive in development environments where they can use functionally rich and powerful command-line tools (i.e., having a bash/zsh-like shell available to them with all the regular tools) – it’s because the tools are so good!

This is anecdotal, of course, but my primary dev environment is Fedora, even though I’m a PC-focused game developer creating a cross-platform game-engine and the largest target platform is Windows, by some margin.

[Aside: this platform preference is not merely because Linux is much more stable, tractable and “metrology-friendly” than Windows (especially when debugging weird GPU-related bugs for low-overhead stateless APIs like Vulkan), but also because AI-assisted development works much more seamlessly on Linux than on Windows (at least in my experience). I have numerous Powershell-related sagas of frustration while using AI, and WSL2 is not a viable alternative either as (a) non-headless/actual game-window-launch related integration tests fail after builds (as there is no Wayland-compliant compositor available out of the box and I can’t be bothered to set one up, especially if the point of testing on Windows is to use OS-native windowing/compositing) and (b) it’s slow.]

Anyway, long story short, the tools matter immensely. And these, whether it’s the Linux-based command-line tools or complex software like Blender (or Unreal / Unity, for that matter), are what I term as the “capital goods of software engineering”.

Much in the same way machine tools like lathes and milling machines help us create tiny parts that go into everything from cars to dishwashers to turbines and spacecraft, the greps and the seds are the hammer and the wrench and the more complex stuff like gcc, ffmpeg, Python or Blender are the CNC machining centres that help us bend, shape and compose the bulbous metallic mass of skeleton code and markdown files into software Maseratis (well, at least that’s the ideal! – for better or worse, vibe coding exists).

Now I know that “devtools” literally has the word “tool” in it but I feel these being first-class members of the capital goods club is an under-emphasised fact.

Capital goods (tools) are harder to create than consumption goods (content)

Let’s go back again to physical capital goods. Think about the complex lithography machines that help imprint molecule-sized circuit patterns on photoresist-coated silicon wafers to help create chips that drive all the smartphones and tablets we hold in our hands today.

Those are capital goods as well: hand-crafted to be functional, precise and fast. They have lenses and mirrors ground and shaped to atomic precision, gargantuan but ridiculously pure EUV light-sources, complex mechanical actuators, vibration isolators and well, software worth millions of lines of code to boot. The most advanced capital goods are often more complex and harder to create than the consumption goods they produce (ASML has only produced ~300 EUV lithography tools so far since their deployment for mass-scale production around 2019 while billions of chips at 7nm or below have shipped since then).

Indeed, how technologically “advanced” an economy is often indicated by the quality and complexity of the capital goods they are able to make: the US, Germany, Japan and, increasingly, China lead the way here. And the companies and countries that create these capital goods guard the engineering behind them very closely, even to the extent of putting export controls in place to thwart competitors from reverse-engineering these technologies.

Contrast this with the world of software-engineering where a lot of the tools / “capital goods” are readily available as free and open-source. And just like their physical counterparts, they have also been hand-crafted to be functional, precise and fast (of course, now a lot of their development is perhaps being done with AI-assistance).

So far, AI has become exceptionally good at learning about processes that can help use, write and test software using tools. It can use frameworks like Django to create very functional websites or and use containers like podman to deploy microservices on the cloud. It can use browsers, their JS engines, HTML and three.js to create great-looking one-shot games rivalling early NES titles. It can come up with a way to use Blender or Photoshop reliably to produce great-looking content (but perhaps not replicate the creativity of the person using those tools, which is why those jobs will eventually survive).

Are AI models good at software “capital-goods” design?

The key question then is, can AI models actually spontaneously come up with the actual tools of software engineering, as part of the solution when tackling a larger problem?

The superficial evidence definitely points that way: for example, I routinely see AI models come up with tiny, reusable python scripts to validate specific parts of their solutions to larger problems – they can definitely be classified as “tools” – much in the same way that bending a piece of wire to become a hook or a paperclip forms a tool. (The word “spontaneously” is doing some heavy-lifting here, I acknowledge, as this apparent spontaneity is perhaps part of the pre-training dataset or RL-based finetuning / feedback that the model went through. For this discussion, I’d call an unprompted mini-tool creation as spontaneous).

But does this tool-creation ability extend to (a) larger tools (like creating Python or gcc itself from scratch, with no prior examples to draw on) and (b) novel problems?

On both these counts, the evidence is mixed: sure we have seen entire bootable kernels and compilers being made by long-running AI agent-swarms and we have seen LLMs solve open problems in mathematics.

But, in the case of compilers/kernels, there is a ton of prior open-source examples these models can rely on. Would LLMs have come up with something like C if all they had was assembly language to go upon or C++ if C had already been invented? That would be amazing.

That brings me to open problems in mathematics: a lot of evidence suggests that LLMs are ridiculously good at taking concepts from various branches of mathematics and combining them with in-depth published literature of the branch of mathematics a problem belongs to, to come up with novel proofs, with tools like Lean being an important part of the solution framework. Would LLMs have come up with something like Lean themselves or, more importantly, the problems themselves – like the Langlands program? Would LLMs have come up with calculus had they the knowledge then available to Newton, Leibnitz and their contemporaries?

[Another aside: to test out this idea of whether an LLM could come up with some novel framework, I gave Gemini 3.8 Flash, GPT 5.6 Sol and Claude Opus 5 (arghh, my wallet!) the task of creating an inter-process communication (IPC) framework from scratch which is quite common for server-client models in game engines and editors. Instead of something new – and despite prompting them to be novel – all of them settled on UDS and named pipes. This is not a criticism of the solution as yes, UDS/named pipes would be the standard way to do it but none of them even tried exploring other stuff. E.g., if one were building their own game console, things like eBPF could be on the table]

Software capital-goods design intrinsically requires innovation

Tying this back to how humans have created the most advanced physical capital goods: even though the engineering processes and recipes might be proprietary, the physics, mathematics and chemistry that underlies it all is open to everyone. Engineers use the physical laws and mathematics to compose mini-solutions / tools together to come up with bigger ones: we learnt how to make precise gears over centuries, precise optics over a similar period, learnt how to use quantum mechanics to come up with super-sensitive yet usable photoresists and the mathematics, built brick-by-brick, to keep it all tractable.

The mini-tools and solutions might feel obvious or simple yet their unique compositions and combinations are definitely anything but! Figuring those out comprises genuine innovation.

There is another dimension to innovation: quality.

You can vibe-code a consumption-good in many cases: perhaps an external-facing high-profile corporate website is still best done by a professional but an internal dashboard may be vibe-coded (and may not require, say, an SAP consultant).

But what makes vibe-coding possible in the first place is the reliability of the tools that actually form the bones of said vibe-coded app: let’s say you have a vibe-coded UI – it may rely on React Native which relies on Skia which relies on Metal/Vulkan and so on (turtles all the way down!). These frameworks are the capital goods which LLMs then leverage. And each of these capital-good layers intrinsically has to be more reliable and performant than the one it sits atop with the quality bar paramount at the kernel level.

And quality requires trade-offs, trade-offs often require subjective judgement which, in turn, requires creativity to begin with.

The future

A single human may be outclassed in speed, base knowledge and correlational capabilities compared to even very small LLMs (say 30B parameters or so). But even the best and biggest LLMs are nowhere close to the average human in terms of pure (relevant) context retention, rejection of the patently impossible and (yet) imagining the obviously fantastical.

None of it is to say that AI models may not get there someday but it doesn’t feel like the simple architecture of searching for the best vectors in a high-dimensional space that fit local context (attention) yet survive the test of world-knowledge (FFNs) are the ultimate solution to that problem (it most likely will be a key component of that solution – the universal applicability of transformers has put that question to rest).

For now, we have created an amazing capital good in the form of LLMs that will help us create more software capital goods going forward. But it seems the purpose and design of those tools will still largely fall on us to tackle.

Coming back full circle to the Arena AI video mentioned at the beginning of this post, the open-world FPV game looks very similar to Halo, released 25 years ago!


About
I like exploring ideas in gamedev, technology, science, economics and (sometimes) culture but I am by no means an expert in any of these. I also like to write from time to time. This blog is an intersection of the two.

I can also be found on Bluesky and Threads. I also like to take photographs, some of which can be found on Instagram.

Disclaimer
All views expressed here are my own and, in no way, are to be construed as recommendations to engage in any securities transactions.

© All rights reserved. Akshay Mishra

Leave a Reply

Discover more from Sum over paths

Subscribe now to keep reading and get access to the full archive.

Continue reading