Oh my god, and the author even supplied a "proof"[0] visual diff harness... that it replicates the original game pixel for pixel.
Just the cherry on top of great demonstration of our collective new superpower: asking computers to do something we can describe how to do, but would (probably) never take the time to do ourselves.
tbh I thought this was a commonly used technique even prior to LLMs? I know I've been using it extensively myself, but I was inspired by Dolphin's extensive visual CI system.
If it is common in the world of video game porting, that just shows my ignorance. I'm familiar with visual diffs in CI for e.g. web development (comparing a static component), but to do that to compare frames over time in a video game/3D environment is new to me.
There are so many more degrees of freedom, which I can see Claude handled... mipmaps, subtle differences in lighting/positioning/compositing etc.
Even then, visual diffs were pretty flakey for web development, because one's OS and browser choice would slightly alter the exact pixels blitted to the screen. At least this was the case for the tests that would simply match pixels instead of computing a sort of visual hash.
It's also partly why some people preferred snapshot tests that compared the DOM tree instead, though that was brittle in other ways (e.g. tests would break if an application's frontend used a major UI library and an update to the library permuted the order of classes in some part of the HTML).
For web UI tests this is mainly solved, at least when using Playwright. It allows setting thresholds, percentages and some other config items to allow some small differences in pixels. https://playwright.dev/docs/test-snapshots#options
I'm getting a 1997 PC game to run on modern hardware and fixing bugs and upgrading graphics as I go, and the amount of quality support tooling Claude is producing along the way is impressive. Fully headless in-memory execution (which, among other things, is used by it for per-pixel diffs too), logic VM devompiler and visualizer, asset explorer, CRT simulator... I just say what I'd like to see, and Claude does 120% job on it each time.
I do that on my renderer. each commit when it's ready to be merged gets a class of visual diffs. Dolphin's way was a major inspiration how to structure it, but it was _the way_ in rendering way before it.
The visual diff harness was probably how they got rid of a lot of visual bugs, just tell the LLM to keep going until the pixels match exactly as the verification criteria
How long before the same thing is done to like, banking back ends? Wallstreet proprietary software? Amazons logistics and distribution systems?
It seems like we might be weeks/days/hours before a situation where someone back engineers and spoofs a system so pivotal to modern human society that the plug needs to be pulled.
Like I've said many times: LLMs are useful for this (i.e. porting software from one programming language to another).
Porting software is painstaking grunt work which still takes a moderate amount of intelligence. It's therefore extremely expensive to port say, COBOL banking software running on mainframes, to another language like Java. That's why a lot of COBOL software is still in use. I expect this to die out in the coming years as many of these systems will finally be ported to another language (could be Rust or any other language).
> spoofs a system so pivotal to modern human society that the plug needs to be pulled.
why would a recreated system be detrimental?
If currently there's a monopoly on a software, this AI recreation is a good outcome to poke holes in that monopoly. It's only bad if you are financially invested in said monopoly, and this would be a minority compared to the amount of benefits that society at large could obtain.
This is basically a digital era anarchist view - the problem you're overlooking is that a lot of critical infrastructure we rely on runs on systems that are considered security through obscurity. Software most people probably wouldn't even know or care that it exists. If you can break the trust of vendors by being able to spoof their proprietary platforms, a lot of the highly efficient networked systems becomes vulnerable to injection and abuse if you can't trust whos making calls to it.
In the past you'd need nation state actors with considerable budgets to do this kind of thing, and we're on a trajectory that could see any kid in his bedroom could do it.
revealing that security thru obscurity is broken can only lead to a better future, even if in the intermediate one there are lots of breakages. It's suffering that needs to happen, and better sooner rather than later imho.
And i assume you don't truly mean spoof as in man-in-the-middling someone - i assume you mean the end user knows they are using an alternate system and are not being defrauded. Like using a photoshop replacement.
> And i assume you don't truly mean spoof as in man-in-the-middling someone - i assume you mean the end user knows they are using an alternate system and are not being defrauded.
I am, actually, talking about MITM attacks that are much further in scope than just defrauding some people using their banking app. I work in resources and operate HMI systems that are networked, but not exactly the bleeding edge of modern software development. If you had an ability to decompile it and recompile your own version you could start sending instructions to infrastructure all over the country - the only thing stopping you is the keys, which if you're intent on hacking someone you'd have the means to obtain anyway.
I can see a lot broader attack vectors than just stealing peoples money. It's the erosion of trust in the api calls themselves.
If someone recreates Amazon's logistics and distribution systems they could try to compete with Amazon? But they'd also need the connections, distributors, transportation, etc. same with banking software, you need capital to be a bank not just software, and if they have the capital then the technology is working we intended making it easier to make new things and innovate, or at least just compete?
No, I am not talking about "taking over" companies and trying to emulate them and do business yourself. You just need to be able to break trust in the api calls and no one knows if a purchase order or transaction is legitimate.
Obviously you need to have access to the keys, BUT I don't see this as a dealbreaker anymore because you just get your agents to go and find them.
I think you’re saying ‘being able to do this means the opportunity for more fraud, by producing fake XYZ as proof’.
Photoshop has been around for around 35 years, fraud has always been an issue. There are plenty of reports of people selling things via Facebook marketplace and the ‘buyer’ showing them sending a payment on a fake baking app. Fraud will always exist and I don’t think tech will make it worse, everyone needs to be more cautious and tells friends and family to be the same.
I'm talking about cloning hmi platforms to send fake instructions to offshore oil platform valve bodies or insert false market trades to collapse companies.
Ah. So in that instance are those platforms not validating that the things submitting information are correct and true.
A bit like a utility company needing to do manual reads every now and then to ensure they are getting a correct signal.
Or they can just take your money and not send out anything.
Alternatively, you just act as a middleman drop shipper and slightly raise the price more than Amazon’s and skim the difference. It might be a while before they find out.
That’s an insecure design. The way we do it here is that you install an app and register your register number and payment card in it. Then when you drive in and out from the parking lot your license plate is scanned and you’re automatically charged. There’s only two providers so it’s not a huge hassle, if there was a single app per garage it would not really work from UX perspective.
I’m already seeing videos of people who have used LMs to reverse engineer and clean room reimplement entire video games. I estimate this shit is minutes away from being shut down, because as we’ve all seen companies stealing is OK, but individuals stealing is heinous and a crime.
I mean some people (me included) have been begging society to pull that plug since over 10 years now.
The plug being "the cloud" and "hooking everything up to the same internet".
These confusion attacks can only confuse people, because critical systems can exist in the same space where entertainment systems and all other categories of systems live.
This was wrong even before LLMs.
Adding to what you both said. I run pixel checks against live websites in a real browser, and I stopped doing full-page screenshots early on: scrollbars, lazy-loaded images and font timing shift pixels between runs like clockwork.
What works for me is diffing small stable regions, one component or one flow at a time, with a small per-pixel tolerance. I also keep a DOM-level assertion in front of the visual check, so when something fails I already know what changed structurally before I compare two images.
For the game port it's genuinely harder, the renderer decides the pixels, not the test. But the region idea transfers: lock the viewport and only diff what's stable.
Actually Claude just did something similar for me, as I'm working on something else with Quake.
On its own it decided to do demo playbacks and take periodic snapshots, and compare them pixel by pixel if the PNGs differ.
I was also working on a web-based port but it saw the original Quake code was not modified so it compiled a native version on its own to use for this.
It had to apply a small patch to make the game completely deterministic, but it figured that out on its own by reading the code.
It then asked me to record a demo with various elements, say an explosion or being under water, and visually verified that the screenshots had those elements present.
So now I have a solid set of tests to verify against.
Opus 5.5 Medium. Used at most 1% of the weekly limit of my $20 plan.
Quake is to some extend a reference implementation of a modern FPS-type game. Because it is open source and because it is well written.
Now I'm no hero in C (doing slightly better in C++). But I have quite some experience with Rust. I rather read Quake in Rust than in C. So I'm happy with the port.
Also, you talk down on the effort but porting over a codebase this size is no small feat even with the help of LLMs. I think the "author" (human creator) learned a thing or two along the way.
Mate its been a couple of days, I would think you'd be mad to expect perfection from an LLM reverse engineering job straight away.
As a proof of concept however it shows that we are at the stage where any malicious actor has a very low barrier to entry to cause large scale corporate espionage etc.
Do you mean that bad actors can reverse engineer photoshop and hack its users or that bad actors can whip out plausible photoshop clone with vulnerabilities and backdoors and hack users with that?
I mean that proprietary software can be reverse engineered, there are many ways that this can affect vendors and end users. Both your examples could be some attack vectors.
Think more broadly than Adobe and any company that relies on a software product could have their business model destroyed overnight. More broadly than that, you could rupture trust between suppliers and vendors if they don't know who's really making requests from a proprietary platform that's been compromised unknowingly.
But thats just good old hacking, you can already see where software connects and try to see if you can get in.
I'm much more worried about AI made software that is developed so quickly that nobody knows what it does, pdfcraft was published a week ago and it has had 5 releases since then.
Oh I obviously know that this kind of thing is already achievable, but it's the barrier to entry is getting substantially lower... Instead of needing a nation state actor that's got funding and resources behind them, we're getting close to the point that some degenerate kid in their bedroom could say "hey claude find me the api keys and services behind nasdaq and build me an interface that spoofs transactions to XYZ". Or send commands to traffic light controllers. Or every third transaction from XYZ bank on every second day has one cent transferred to this other account.
Not just api calls but the business logic too.
Laymens thoughts there, but in essence this is my concern. People just looking at any old service in the world and going "hey claude build me a hook into this" and the models are getting close enough to being able to find and decompile the right moving parts on its own.
They're RL'd to oblivion into "solving tasks". There's a lot of things that they'll refuse to do if you tell them, but will do if they decide it's the shortest path to "solving the task".
I hope the future is brighter because the present is bleak.
Models got better at doing things the users are entirely within their moral and legal rights to do, without punting or forcing to argue the point.
In my little game restoration project, Claude just holds me to the commonly accepted standards and makes sure none of the original assets ever make it to the git repo its working in. I only once had to assert that occasional game screenshot to illustrate some doc is acceptable - again, it didn't argue the point, just said something about abundance of caution.
The "clean room" approach here is using public documentation to assemble a roadmap/guidance and computer use to get any other behavioral output and visuals from the software you wanna clone. The agent doing that will just do something which looks completely innocent to it (e.g. create a rectangle on the canvas and then rotate it).
The other agent then implements the documented journey.
The endavour doesn't have much value at all except as an experiment how well translating C code to Rust via LLM works (but you don't need an LLM for that either, C2Rust was already a thing).
Also, the C version of Quake compiled to WebAssembly and running in browsers is just as "safe".
But C isn't as nice to work with. Maybe OP just thought it was fun to see it take shape and wanted to share it. Why do any sort of little side project like this at all? Just seems like everyone is judging this too harshly. If an LLM is reasonably good at this sort of thing, and you think it would be cool to see Quake written in Rust running in the browser, I say just do it and have fun and don't be shy about doing show and tell about it.
All commits in the project are from Claude. What's fun or to be learned from that?
It's of course everybody's personal choice, and it's nice for an experiment to figure out what Claude can do. But why go on HN and brag about a project entirely created by Claude when literally everybody else can do the same thing without lifting a finger?
I see the result of an LLM port only as a baseline. I would expect any serious developer to rewrite and refactor the code beyond a verbatim translation and to make it more maintainable, readable and robust. In short, to turn it into idiomatic Rust.
If you're not going to do that don't post your results online, just keep it to yourself. Anyone could've done this. Even people without any programming skills.
The definition of "serious" developer is changing. An LLM can already produce maintainable, readable, robust code.
It's getting to the point that LLM produced code is indistinguishable from even the best handwritten code.
There's no value in gatekeeping,the code side of software is becoming increasingly democratised, and the barrier to entry is getting closer to no barrier at all.
> If you're not going to do that don't post your results online, just keep it to yourself
...I mean...I just don't really know how to respond to that. Who the heck cares if someone posts this online? If you don't like it, just move on. Sheesh. Lookout everyone, this guy doesn't approve of your side project.
Does budding a cabinet in a weekend counts as any sort of project, given that you used factory made wooden planks and power tools, instead of starting by slaughtering your horse with your bare hands, so you can make a knife from its bones, and use that to make a saw from its sinews and hair, and use that to fell a tree, and...?
Those who remember the time before a new technology will always have a different world view than those raised with it. Some of us will prefer it, some of us won’t. Eventually no one will remain who remembers. That’s just how humanity works.
It does have various improvements over the original version. And I don't agree with the notion that "all work that relies on LLMs requires no effort". If this were the case, we would have plenty of complete Rust ports of Quake. It's open-source, allegedly super easy to port. But as far as I know, we did not.
What this exploits is the mental shortcut of "rust = good", which might be a good thing, as that was always wrong. But now it is being pushed to its breaking point so that that idea will eventually collapse.
Accelerationalism on a micro scale, basically.
AI keeps breaking things that were broken before like this constantly. It's the great cleanup of old bullshit. (Unfortunately through even more bullshit, but at least there is a silver lining)
Author here (someone else submitted it, thx). There's a video as well, a six-minute video on how it works and the various QoL that we added (made by AI with human assistance):
This took about 4 months as a hobby project; I don't know how much of that was work and how much was "waiting for the next model", probably a mix of both.
Since this ended up on Hacker News, I'll say it: I find the discourse around LLMs strangely sad. They're annoying, yes, and they make mistakes every day, every hour. But they're also amazing, and working with them has been one of the most rewarding things I've done in years.
Historically things that make people more productive have tended to make us richer, not poorer. And it's a virtue to change your mind as the evidence comes in.
We played with Lego all our lives; now the Lego understands us and helps us build. I think that's something to be glad about.
You're putting the cart before the horse in various ways. Saying "it's a virtue to change your mind" implies you're already at the correct solution, and everyone who disagrees with you is wrong.
I would never have imagined a phone web browser capable of this in 1996 between playing Quake and testing out the hot new JavaScript powered mouse rollover image effects in Netscape Navigator 2.0
wow, that must feel quite surreal. i’ve been using the internet since 2001-02 and became a web dev in 2006-08 (discovered what js was), and even I find this impressive & fun (especially that the code isn’t written by a person). But to see both JS and Quake as fresh new things to here in 30 years would feel crazy.
Quake ran reasonably well on a 486DX100; what’s the interesting part here? We already know LLMs can handle porting, and it seems to have become a trend among attention-seekers to convert just about anything to Rust.
I think 'standing on the shoulders of giants' is the phrase for something like this. It's the confluence of browser rendering, WASM, Rust, and LLMs. For me it's less a demo of what AI can do, and more a showcase of the human effort from the past few decades on the parts that needed to fall into place for an LLM to come in at the (relatively speaking) last second and claim a win. Sure, an LLM did the port from C to Rust, but think of all the things needed for it to all work. That's pretty damn amazing, and it wasn't done with AI.
This may sound funny but I feel games would lose a lot of fun if they were all written in rust and had classes of bugs just not available to them. For better or for worse quirks and bugs in games have shaped how people approach games, and also have given games charm for decades.
Rust doesn't check for integer overflow in release builds by default (ignoring the explicitly-checked methods, of course). At least when building with Cargo whoever builds the binary sets overflow behavior.
Fortunately for gamers, Rust doesn't do anything to stop physics engines from going haywire or preventing players from clipping out of bounds. A Mario 64 written in Rust still has parallel universes (well, assuming that you carefully translated the out-of-range float-to-short cast as having modulo semantics, an operation which doesn't have any defined semantics in C).
C doesn't bounds check arrays. A C compiler is perfectly capable of producing the same machine code as an assembler when given a loop that writes bytes to an array and then keeps writing beyond the space allocated for the array.
Has anyone done counterstrike clone in the browser? have friends join in and throw frag grenades. The open map ecosystem provides great free textures/artwork already.
If someone is looking for the next srp - might I suggest CS 1.6? it should run find one any modern browser.
‘Safe rust’ just being used to mean ‘no unsafe’ is kind of a shallow understanding of the language. You can write code without unsafe that is not really ideal at all eg abusing vector/slice indexing to create a kind of interior mutability that the borrow checker is blind to. Which as a C port I’m going to guess it probably ends up doing
Almost hard to believe that someone using an LLM to mindlessly migrate software from one language to another for no practical reason at all would have a shallow understanding of those languages...
I remember seeing the QuakeC line of code that halved self-damage from rockets. It enabled rocket jumping, but it also made the rocket launcher a much, much better close quarters deathmatch weapon than in Doom, so I thought it was a bit lame.
Vibe coders need to stop trying to use the browser as a platform for complex games. It's never going to work out and will always lead to a slow, unplayable, piece of shit. If you want to make a game then just make it a desktop program... There's a reason why literally every modern game is built in this way.
These aren’t JavaScript apps, they compile to WASM which actually has really good performance. Probably not as good as native assembly but for an old game it’s easily more than enough.
Adding support for larger maps and colored lightmaps, for example. Quake still has an active modding community, and many of the most impressive maps require a source port with increased engine limits. The original engine just can't handle their scale and complexity.
Luckily your tokens are equally accepted as legal tender to the code machine, so if you want it, you can have it.
It a new world. You can literally maintain your own fork with the features you want.
Even the AI models are not gate-kept, soon. In 2030 you'll be able to affordably buy a GPU that can run an open weight model equal to the best that 2026 had to offer. I.e. projects like this.
Just the cherry on top of great demonstration of our collective new superpower: asking computers to do something we can describe how to do, but would (probably) never take the time to do ourselves.
[0] https://github.com/terrapapagalli1516/quake-srp/tree/main/or...
reply