This looks pretty decent actually. Sure, you could consider it a frontend/SDK for bubblewrap/seatbelt/processcontainer; but setting em up consistently is far from trivial; and hand rolling is a really bad idea (speaking from experience).
I like the ‘learning’ mode for figuring out what perms/config a runtime needs, the MIT license, the clear optional telemetry disclosures, and somewhat light and still readable documentation.
Regardless of your views on Microsoft, this looks quite useful; serves a clear purpose, and from a quick glance, looks like a high quality project even if it’s just the first version.
I agree! I can't comment on the design of the SDKs but from the looks of it this is a great option to integrate into agent harnesses/pipelines.
The "audit" and "debug" mode are specially useful. I use bubblewrap and using a new harness with it is usually a couple of rounds of wack-a-mole with strace to figure out all the harnesses dependencies.
Do any of these sandboxing solutions have a dynamic component to them that lets you grant permissions, starting with a minimal sandbox and asynchronously adding permissions as they become necessary? Harnesses try to do this when accessing non-project folders, but it's not always strictly enforced and generally not revocable. Harnesses also block agent execution until a decision is made, which requires constant monitoring to ensure progress can happen when the agent could easily proceed with an alternative method right away.
I like the idea of a minimal sandbox that protects against accidental `rm -rf` and against personal data leakage, but such a setup then often gets in the way of the specific task to be done. Ideally the sandbox would be able to aggregate blocked accesses and then expose them in an external TUI dashboard, where I can then enable access (without blocking any running agent on this, since that's prone to "press okay" fatigue).
I think this may be a false choice because of the work pattern you might be used to. Assuming here, If you work in interactive sessions where it's open ended there is a boundary where you have done enough research/prototyping and you need to move to implementation and the permission scope has to change now. I think the realization that you might have is that if you're doing this then it's probably best to separate the automated AFK part from your initial research part.
Even for pure research/prototyping, you quickly run into the problem that your sandbox is either prohibitively minimal or overly permissive. Depending on the exact task, you may need GPU access, Docker/Nix socket access, ability to ptrace processes (gdb), run webfetches, etc. If I define a "research" sandbox profile to allow all of these, I might as well not have any sandbox at all.
If I understand you correctly https://nono.sh/ might go into that direction. It can add permission after the sandboxed command is terminated based on which blocks occurred.
Its not life though as you seem to describe
Very interesting concept to make permissions specific to individual shell commands though. Certainly good that people are experimenting with these approaches, hopefully ideas will eventually converge so we don't need to know like a 100 different sandbox projects :)
I have faced this exact dilemma as well. It's not always clear what permissions are needed in advance for my pi sessions and it's child sub sessions. A simple example is when child sessions do a task they locally want to fire random docker commands to learn the state of my local docker devstack.
Sounds OpenShell locks filesystem access on sandbox creation - arguably the most important isolation feature, at least for my use cases. Architecturally it looks right though!
A little off topic, maybe, but I've been having great luck with wasmtime and wasm32-wasip3 for writing sandboxed plugins. The tooling is pretty nice when you write plugins in rust, but I don't know what it looks like for other languages right now.
wasip3 is not stable yet, but it has a lot of nice changes (compared to wasip2) for integrating with async code
That site thinks this file has 2.9k sloc and doesn't seem to parse rust comments. In reality, there's only 1,465 sloc; and 635 loc of tests.
Definitely nowhere near 350k sloc.
--
As for your sandbox run: it's a single-contributor project, seems to have only have basic smoke tests, and has a few major/critical security issues:
* _generate_seccomp_filter compares newline-deliminated syscalls, against a multi-line blocklist, meaning the entire function doesn't block anything and is essentially a no-op.
* Main script invokes working directory's .env as shellcode, before switching into restricted filesystems and dropping capabilities. Attacker-controlled .env can run shellcode with full privileges.
* Lots of race conditions which I haven't verified, but doesn't really matter.
I'd make PRs, but I don't think it's a good idea to try and DIY a sandboxing system in bash with minimal SLOC as the target in the first place. I'm also slightly concerned that most of your comments on HN seem to be promoting this repo?
Thanks, I see there's a slight (~50%?) overestimation there, but then again, even unit tests and comments in a target programming language count as syntactically correct code that needs to be evaluated and reasoned upon. I'm not that familiar with Rust's runtime introspection features, but in languages like Python, even the comments can directly affect code (e.g. `Foo.__doc__ = Bar.__doc__ + SOME_ANNEX`).
> _generate_seccomp_filter ... the entire function doesn't block anything
Many thanks! I've applied a fix—it's a single line added. The missing test is pending a runnable that invokes one of the forbidden syscalls. As I have no qualms about force-pushing around a repo that nobody forks, happy to credit you(r LLM) proper!
> Main script invokes working directory's .env as shellcode
The sandboxed process can't overwrite existing .env files [1], but it could create a new $PWD/.env file, hoping to "escape" at next sandbox execution. That's a valid concern I'll have to think about some more.
I sometimes experience "Slirp not ready in time" [2], but it's due to a so far unexplained upstream issue [3]. I you have time/tokens to spare, I'd appreciate those PRs and further similar feedback!
Don't know whether it's a good idea. It sure has got its issues. But even as the SLOC count and the number of bugs metrics are proved correlated in literature [4], min SLOC is not the primary target—a reasonably graspable and stable composition of few dependencies is. Whereas overreliance on third parties nowadays often ends with a rug pull one way or another. We simply can't count on this "MXC" (...) to be maintainable/non-archived even a year from now, just when I'd get it all properly integrated and set up.
Oh, I certainly wouldn't like to limit myself to promoting just this repo! ^D^ HN is a good venue, lots of smart people around! I see everyone shilling their own sh** all the time. Often in green usernames. :shrug:
Not that I trust either of these, but at least I can sit down and read 500 lines of Bash. I can't read hundreds of thousands of lines of Rust, and although in the past I could assume that someone at Microsoft would have reviewed it all, these days I'm not even sure of that.
The main usecase for an average developer is preventing confused agents making mistakes like removing sensitive folders, resetting git branches or using API tokens they shouldn't be using
Why is everyone making their own code execution agent runtime engines I have an entire project built on top of openshell already, why not first come up with a sandboxing policy design, like unix did, and then build on top of that.
Currently all project do tend to agree on what and how they work but certain things being different makes porting tedius, if all of them have a bare minimum subset common amongst them it would be much easier to switch, and validate security surface area.
I feel like there are more vulnerabilities in this vibe coded slop sandboxes, and it's more likely everyone one of us trusting them to build projects around them will shoot our foot off once a cve is hit in one that's common in all of them but since they are all slop copies someone will have to figure out how they apply to all others and then manually fix it properly, and if one of them makes a CVE public it will leave dozens of these runtimes open to exploits.
I wish the best to my future self with regards to security I feel like we are completely screwed. Since we can no longer depend on upstream for security.
"Im really fkin worried we're all building the same thing"
Everyone has been building a harness/sandbox the last 6 months. Ive seen dozens and dozens shared in discords.
Even companies are totally stuck focused on the same paradigms.
The previous iteration of this was RAG/Chat interfaces. See PewDiePie's project. Last month it was briefly everyone building the same classifier.
Peter Thiel, gave a lecture about this same phenomenon 15 years ago likely because he observed the same things going on during other hype cycles. Everyone building the same things. Its called something like "Dont build the obvious thing"
This is why Im moving towards hardware for personal projects, it forces me to be much more creative and think outside the "How can I make something AI adjacent/powered" trap thats so easy to fall into in pure software right now.
My guess is that Microsoft thinks this can be deployed with standard, company-wide policy across platforms (mostly) with their IT management tools which poses a unique advantage.
In reality, however, knowing how much difference there is between OSes, how tricky it is to configure these things to make them actually useful, and how bad Microsoft products are, I'm not enthusiastic about this project -- there are so many others on the market already, and I'll wait to see if this gains traction.
(Notice that on MacOS it only supports seatbelt? That's not nearly the same as microvm.)
The problem with OS level sandboxes, and the reason why WebAssembly's being explored in this space (and Electron is so popular), is that relying on OS/hardware features means your TAM shrinks to a fraction of total, and it's historically well known you set yourself up to lose.
History is littered with tons of super cool OS features that didn't manage to gather enough market share and ended up as cool futures, and fodder for 'we invented the future 20 years ago' style articles.
e.g. this one puts multi-platform support as a high requirement, a requirement that OpenShell doesn't fulfil (and likely won't given it's architecture/goals).
That’s exactly it.
There should be some permission logic for delegating.
But as always, there are rivals trying to set their tone on what’s the standard. We all wish there was one unified agreed concept that will work but I guess the most common one will eventually survive.
Just as Microsoft in a sense embraces Linux with WSL and also Apple has their virtualization framework.
I hope we’ll eventually get unified model management system to include also permissions designed properly
They have a sandbox escape in there. Likewise capability ordering is wrong. Exactly what you should expect from Microsoft.
Since this apparently wraps bubblewrap (another incompetent action on behalf of Microslop), did a quick sweep of that codebase too. Setuid is wrong, capability dropping is wrong, bubblewrap does _not_ protect against compromised/vulnerable kernels (and it should fyi), wrote up a full bubblewrap/mxc sandbox escape too.
Could you link to issue reports to back that up (or maybe file them if these are novel findings)?
If that’s too much of an ask, at least reference the code you found for things like “mxc sandbox escape” or “bubblewrap setuid is wrong”. Those claims require evidence.
Asked Grok the difference between mxc and flatpak:
"So MXC is a cross-platform “what may this workload touch?” layer aimed at agents. Flatpak is a Linux app format whose sandbox happens to share a backend with MXC on Linux."
They have this for Windows and Linux, but it's sadly missing for macOS - see the support table here: https://github.com/microsoft/mxc/blob/main/docs/backends/sea...
Things macOS is missing include "Allow/deny by hostname" and "Allow/deny by IP, CIDR, port, or protocol".
The rest all looks great, and if you are on Linux or Windows those restrictions don't apply.
I guess this is the universal challenge of building an abstraction layer over multiple different technologies.
reply