Hacker Newsnew | past | comments | ask | show | jobs | submitlogin
AI-ready biological data: $1.8B global commitment (biohub.org)
133 points by ray__ 16 hours ago | hide | past | favorite | 19 comments
 help



TFA’s missing info in between its PR pairing of the unknown biohub with the (much more well known) NIH and DoE is that biohub is a Zuckerberg initiative. So just imagine all your biological data paired inextricably tied on an opt-out-but-cant-really basis with your Facebook & Insta account.

Yea I was thinking about this too. Then I thought about how this will be used elsewhere. Predictive models for mental health conditions, addictions, diseases, etc. strong data to try to exploit for marketing. Then eventually health and life insurance companies will be deregulated enough or be willing to pay the fines to squeeze some more capital that way. And of course one day if we are lucky enough eugenics efforts and specifically targeted bioweapons. Yay

I’ve been toying with the idea of a SETI@Home successor that uses leftover subscription credits from the (currently) subsidized AI providers to help contribute to a collective goal like this.

I wonder if there will be any opportunities to combine resources collectively like that again in the future.

I built a small prototype that would let people in less-developed countries get AI responses from it for free. But it’d be super neat to work together and put our laptops and subscriptions to work on something like this.


After a bit of research, it looks like BOINC[0] and AI Horde[1] are along the lines of what I was building. It appears NVIDIA PAIR[2] has been working on distributed inference for local networks as well.

[0] - https://boinc.berkeley.edu/

[1] - https://aihorde.net/

[2] - https://www.nvidia.com/en-us/ai-on-rtx/personal-ai-router/


Sounds like a great way to get your account suspended… My Claude account was once mysteriously suspended (I wasn’t doing any reselling/sharing and not asking anything sensitive) then mysteriously restored three days later, so I’d say actually violating ToS is likely more risky, given that they seem to not mind suspending false positives at all, and you have zero recourse once you’re suspended. Their customer support is an AI chat bot and even that chat bot page just redirects to the suspended page lol.

You’re right, of course. I never made it out of PoC. But I also implemented something with locally running models to do the same thing. Request comes in, gets forwarded to a machine not currently busy, it serves.

It would be great if we could use the credits we pay for and don’t use on our subsidized subscriptions as well. But, yes, a matter of “when” not “if” for your account getting shut down.


I feel like it would be better to create open synthetic training data for all to benefit from,or something along those lines, as opposed to giving it to those claiming to be in need, as such a system would be exploited and abused in a matter of days.

There used to be folding@home for exactly this; pre-Alphafold days, where you could run protein folding on your machine.

Yes! I used to love watching the visualizations for this. You’ve got the right idea. If it truly is AI processing to help us eventually create this virtual biology, why not our idle machines with models running on them?

To avoid suspension I was thinking of something similar driven by Github issues, pr and a skill people would load and run.

They’ll easily detect that. Remember when they routed subscription users’ sessions to extra billing when a certain word appears in the context? https://github.com/anthropics/claude-code/issues/53262 https://news.ycombinator.com/item?id=47952722

Claude Code is known to snoop on other aspects of its operating environment too, not just the texts sent to them: https://news.ycombinator.com/item?id=48734373


Compute was never the primary bottleneck here. High-throughput wet-lab telemetry and standardized multi-modal ground truth are. Good to see capital flow into actual data acquisition instead of more wrapper layers.

Would you mind to clarify what you mean by standardized multi-modal ground truth?

Paired measurements from the exact same sample under uniform conditions—e.g., spatial transcriptomics, proteomics, and morphology co-registered together. It stops models from fitting to cross-lab batch noise.

We really need to own our own data and allow it to be used for the common good, unfortunatly people just dont get the concept or the value. We are open source and need good open source laws to protect us.

Meanwhile the current administration is actively taking formerly publicly available data sets that are critical for research offline.

We need increasely difficult bio-agi contests.

put that RSI to use here and let it rip

Could be cool

It’s like the Mayo Clinic’s data grant




Consider applying for YC's Winter 2027 batch! Applications are open till November 2.

Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: