Would be a terrible shame if lots of people opted out.
If you are a EU citizen you might also want to write a complaint to privacy@huggingface.co because they are collecting your personally identifiable information in machine-readable form which they are distributing to third parties.
We want to give developers agency over their source code by letting them decide whether or not it should be used to develop and evaluate machine learning models.
fuck them seriously; if you want to do that then don’t steal the repositories in the first place.
Meanwhile at work we just had a training course that specifically said doing “opt out” instead of “opt in” violates the principle of informed consent.
not surprising since the venn diagram of ai bros and rapists is a circle.
Kinda hope it uses my code. It’s so terrible there’s no doubt it will make the resulting code from the model worse even if the impact is miniscule
They cloned a few projects I’ve made that rely heavily on a library I also made, I’m only pulling out the library.
Modern tech companies love using the rapsist’s model of consent.
don’t steal the repositories
How is this remotely stealing?
It’s not, but it may be violation of licenses. And also, if it has personal information on it, that’s probably illegal under the GDPR.
but it may be violation of licenses
They excluded code with non-permissive licenses apparently:
Each file is labelled permissive (at least one permissive license detected, no conflicting non-permissive license), no_license (no licenses detected, or only non-license legal texts such as CLAs), or non_permissive. The permissive allowlist follows the Blue Oak Council list plus licenses categorized as Permissive or Public Domain by ScanCode. Files classified as non_permissive are excluded from both released datasets.
also, if it has personal information on it, that’s probably illegal under the GDPR.
It’s all public so I would be extremely surprised if that were the case.

the method on github is to fork it not steal all the data in an external storage for external purposes without consent of the user.
edit: you can also report their whole account to github here: https://support.github.com/contact/report-abuse?category=report-abuse&report=bigcode-project&report_id=110470554&report_type=user
I’m pretty sure they just cloned the repos. That’s how GitHub is designed to work. Have you never cloned a repo from GitHub?
Yes, but I used it under the licence terms of the original developer.
So have they. They filter by license - see my other comment.
so if you are licenced under MIT and they use your code, they publish your copyright header?
Yes. If they train AI from your code? No, but the legality of that is yet to be settled and definitely leaning towards “it’s fine”.
not for private for profit use from shitty ai companies though…
Unless I’m mistaken, this wasn’t written by the folks that scaped GitHub in the first place, someone just wrote a small tool to semi-automate the process of searching the scraped dats, and submitting a GitHub issue to have it removed.
it’s pretty clear from the language of the site, the fact that it’s on the hugging face domain, and the fact that the github organisation for the “opt out” makes it clear it’s hugging face.
I never really needed any justification, but ever since AI companies just take stuff illegally, and that openly, it has become my justification to just pirate the shit out of everything.
(Excluding indie games of developers I like).
This specifically isn’t illegal though. They’ve only scraped public repos.
That depends on the license. Music is also available on Youtube without authentication requirements but I don’t think you can just download those and do whatever you want with them.
Could be wrong, but the act of downloading music isn’t what gets people in trouble, it’s the uploading. When people torrent both happen, but afaik it’s only the uploading portion that is actually bad.
This depends, but not because of piracy. If you have to bypass DRM to download the music, you may be violating DMCA §1201 (even if you are otherwise allowed to use the music).
Edit: This obviously doesn’t apply to GitHub, but might to YouTube.
When people torrent only a part of the file is being shared by all peers. Technically only the original seeder shared the whole thing and after a few seeders it’s a tiny portion of the file.
I think there’s a difference between the laws for individuals and corporations. I don’t think they’re allowed to download either.
Find me a repo on GitHub with a license that disallows cloning it. I’ll wait.
Any repo without a license disallows reproducing it in any way, which includes cloning it for business purposes.
The only exception is that public repositories can be forked on GitHub as new GitHub repositories. Forking doesn’t give you a license to use or reproduce the code, though.
The best way to opt out is by not using GitHub. Also opts you out of Copilot and a bunch of other stuff.
Already did that some time ago but I am still in the dataset.
Other git hosts are also getting scraped, and have had to implement counters because of it. For example, this is the kind of thing Codeberg shows crawlers. I’ve even seen people who self-host complaining about getting overloaded because of bots scraping their forge
I’ve put Anubis before most of my website, including my forgejo instance. For the projects hosted there, which is not all, I can only hope that that’s enough.
I like to have the visibility and CI of GitHub. But this sucks ass.
A decade ago, if someone asked someone working on an Open Source project if they’d be ok with an AI reading their code and learning from it, they’d most likely say “yeah that sounds really cool!”
Somehow the tech-bros have fucked up AI so much that something that should be really cool seems creepy, lame, and nefarious all at once.
deleted by creator
Making arguments from popular perception is never very strong. AI is cool if you actually think about it - the capability is incredible. Anything that can produce working code was going to have this ambivalent result where execs pushed it way too hard.
Yup, the idea is good. How this this idea is being implemented… not so much.
Simultaneously they have created the part I wanted from Star Trek, while making the worst possible anti star trek a reality.
E.g.
I want to be able to ask a computer about history, art, codeing, well anything. And be able to clarify and question and put together new ideas.
But not by anyone owning that ability or profiting on the labor of others or causing environmental harm.
We got the cool computer but haven’t achieved the post-scarcity part.
I don’t think you can have one without the other.
Post-scarcity isn’t actually possible. Not even in Star Trek this is true. Picard’s family owns a vineyard in France filled with artifacts and antiques. Not everything is fungible. People will desire these non-fungible things. Not everyone that desires these things will be able to have them, because they aren’t fungible. You can’t have a billion people all owning vineyards in France. Some people won’t get everything they want. There will always be scarcity.
Star Trek was made in the 1960’s at the height of the cold war. They didn’t want the show to be about how capitalism was superior to communism, or vice versa. So they side stepped the issue by saying in the future there’s no scarcity, no money, and they live under some ideal future economic model. The show writers don’t know what that ideal economic model is, so it’s deliberately vague and inconsistent. And that’s fine because the show isn’t about economics.
It was wise of the writers of Star Trek to avoid trying to make predictions of that nature. 300 years ago, Wealth of Nations wasn’t yet written, the field of economics didn’t really exist. They believed shiny rocks had intrinsic value. They had some vibes about things like currency devaluation and inflation, but economics was still mostly about acquiring shiney rocks 300 years in the past. So what will economics be like 300 years in the future? None of us know, and certainly writers of a TV show don’t know.
The writers of a TV show saying there will be no scarcity in the future just means they didn’t want to discuss economics. It’s not any kind of prediction about the future. There will always be scarcity.
deleted by creator
There have been periods of human history within some groups that did have no needs scarcity. Shelter, food, water, all there. Replenishing faster than they could consume it.
Star Trek addresses needs scarcity to a large degree, instead of want scarcity. Of course, the TNG makes it a little more clear that they can make anything you want, which who knows if we get there.
A good example would be do you have somewhere to live? Need Scarcity met. Do you want a high floor overlooking the Nebula? Want Scarcity, not always possible.
Of course it could be argued that if you could go multiple times the speed of light in a infinite universe you are likely to find something that you want.
People will desire these non-fungible things.
This is the case where if I were to replicate it, what would be the difference? Do you think if we got to that point we could intellectualize that the same molecules are the same molecules or not?
There have been periods of human history within some groups that did have no needs scarcity. Shelter, food, water, all there. Replenishing faster than they could consume it.
Which periods are these?
I’d say right now we have the necessary resources to provide for everyone’s needs in the world. Providing for basic needs isn’t a scarcity issue, it’s a resource allocation issue. We aren’t prioritizing providing for everyone’s needs before providing for people’s wants.
I’m lucky enough to live in a developed country, I have shelter, a fridge full of food, I turn a tap and I have potable water. It’s all there for me. My needs are taken care of. I still think would be awesome if I owned a private island in a nice climate, with a big house on it and a pier with my own yacht I could sail around when I feel like it. But there’s the problem, how do we determine who gets what they want? If I got the private island, yacht, etc. while you only had your basic needs met, would you consider that to be an equitable society?
Why does Picard get a vineyard in France while someone else has to be satisfied by whatever the replicator provides them?
TNG makes it clear that Picard values original artifacts. There’s a scene where he assumes an artifact someone is giving him is replicated but then is told it’s the original. He basically pisses his pants over it being a much too valuable of a gift. So no, in Star Trek people do place greater value on non-fungible items. You don’t even think of this scene as unusual because of course people think that way.
Starfleet needed more starships to defend themselves from the Dominion. Do you consider that a need or a want? They don’t have an infinite number of starships. Apparently building starships involved some kind of scarce resources.
Because they didn’t have enough starships, a lot of people died. If we consider not being vaporized by a phaser to be a need, the Federation was not able to provide everyone’s needs because of a scarcity of resources limited the number of starships they could produce.
Like I say, Star Trek is incredibly inconsistent about this. If an episode requires scarcity of resources so Nog can trade different things to different people in a fun episode, then those resources are scarce. If they need to use gold pressed latinum, they use currency. If someone asks “how does your economy work?” then suddenly “they don’t use money, they don’t need it because they have everything they want.” And this is fine, if someone 300 years ago were to write some fiction about our time, it would probably be about people more Scrooge McDuck (who had a giant vault of shiny coins) rather than Elon Musk (who has no need for that because that’s not how the economy works). Star Trek can’t have a realistic economy, and it doesn’t need it. We’re just meant to assume it’s broadly equitable, everything is fuzzy beyond that.
This is the case where if I were to replicate it, what would be the difference?
You can’t replicate an infinite amount of land in France. You can’t replicate an infinite number of private islands. People will value these things and they will be scarce. Why would Picard live on a vineyard in France instead of in a holodeck version of a vineyard if there was no value to the real thing?
which periods
Coastal peoples of California, Jamon of Japan, Puget sound Indians, Hawaiians, the Congo, and so on. Places where you have mild climate, hugely abundant food supply and easily worked materials.
I’m lucky enough to live in a developed country, I have shelter, a fridge full of food, I turn a tap and I have potable water.
You do, but I am assuming this all revolves around keeping a job or employing yourself somehow
I have had the discussion just like this “If I got the private island, yacht, etc. while you only had your basic needs met, would you consider that to be an equitable society?” with some people. Interestingly they often say, so? why would i want an island or a yacht? And if I did wouldn’t I just go to the holodeck? Seems like a lot less to take care of.
Which brings us to this:
People will value these things and they will be scarce.
Do they? Or are we conditioned to feel that way?
But yes in the end Star Trek is just a fiction with themes at the core, not exact solutions or even realism when looked at closely.
Either way, my point was I want the computer with knowledge and not have that some how steal, hurt, take away from, get owned by, deprive anyone, or cause environmental damage. And my first thought about that part was post scarcity where no company would own any of this.
I mean they’d have been ok with it because tech bros were the ones automating other people out of jobs and never thought it would come for theirs.
The level of AI we have now was impossible science fiction a decade ago.
A little bit infuriating since huggingface itself requires login to access a large portion of the content on their site
We want to give developers agency over their source code by letting them decide whether or not it should be used to develop and evaluate machine learning models.
crawled directly from GitHub and built to pre-train code LLMs with full-repository context
Repositories that opted out are removed from the dataset before each patch release.
“agency”
Which AI company will not use v1 which has all of the data but will use later patch releases instead which have less data?
Or just merge it
Since they stole my paper on ethics in computer science, maybe the model will learn to act better than its owners
Am I alone in not wanting to put my username into that field? If they don’t have it will they then just decide that it’s now a good time to scrape it? Or are they going to record that it was searched?
Pretty sure all my repos are MIT licensed for the betterment of everyone, but I’m not on the list! So I guess I’m not good enough, or they are failing to follow the attribution clause of it.
My dotfile repo is there and it doesn’t have a licence. Meaning it’s technically not open source. Didn’t stop them
“Oh no, people are using information I put publicly available on the internet for everyone to see!”
Morons. The lot of you.
Edit: I rest my case.
I did it to invite collaboration and connect with other developers with similar interests. FOSS is more about building communities than building software, after all.
I did not anticipate that it could be (legally) used to dismantle the kinds of communities I wanted to build. (I did anticipate that it could be illegally used to that end, but historically that has tended to cause a Streisand Effect, so that risk seemed worth it.)
If people still want to be part of the community, they can.
You’re just coming up with reasons to fit in with the crowd.
Love you too, kitten. ♥️
Sure. Let’s see if they use it for endeavours in the same spirit.
Or are you fine with they using this data for-profit without benefitting the public by also making it open?
(not talking about hf, I know starcoder. just in general)
Maybe you should’ve released your code under a license with that stipulation.
In all honesty, you’re just coming up with reasons to fit in with the crowd.
Give me a crystal ball next time and I’ll do it genius, use your common sense.
Seems fitting the one thinking about “the crowd” the most is exactly the one trying to be the most distant from it. Projecting much are we?
Don’t hurt yourself with those mental gymnastics.
Hey, sure. I’m doing all kinds of gymnastics to think the ones profiting off our public data are in the wrong, I’m just in it to fit in and be in the crowd sure! Whatever makes you feel different. Hope your mirror works someday.
See? Even now you’re still trying to save face.
You need help.
Oh God why are all your replies in the same structure, am I talking to an LLM? I can’t decide whether that would be preferrable or not
tbh i’m thinking this alone isn’t that bad from an archiver/datahoarder perspective
As long as the datasets are open, it is our best hope. I know it doesn’t compensate the people whose work’s copyright and licenses have been violated, but I think it’s the only realistic hope we’ve got of getting out of this informational dystopia with a reasonably intact library of humanity’s knowledge that hasn’t been locked down and/or monetized. The AI scrapers and generators are in the process of burning down the great library of Alexandria that the Internet had become, and we are already starting to feel its loss. We cannot stop the wave of toxic pollution that is spreading through all our digital content now, but the archives from before this apocalypse started will become the most valuable thing humanity has ever produced. This is information war, and we are losing.
deleted by creator
IA is a nonprofit and archives to preserve human history. shitty AI startups do this to monetize the data, and their end goal is to “replace” the people who made that data in the first place.
Thanks to these AI mfs the IA now prevents access to many items because they can be used as training material which fucking sucks
As a developer, you hold the copyright to your code. When you make it open-source, you grant a license to use the code and the resulting program under certain terms.
This is a contract. If you copy my code without following these terms, then that’s theft.
The Internet Archive’s use complies with these terms for all open-source licenses. These AI companies do not. In particular, here’s a quote from the MIT license, which you will find in a similar wording in all open-source licenses:
The above copyright notice and this permission notice shall be included in all copies or substantial portions of the Software.
In effect, what this means, is that when you copy my code, I demand that you also copy the license text along with it, so that anyone else looking at this code knows the permissions I grant and the terms I require.
And now guess what these AI companies are doing. They copy my code and reproduce substantial portions upon a user asking, yet they do not include my license terms. They violate the contract under which they obtained my source code.
I suspect you don’t realize how shit that is, because source code is so abstract.
It’s like spending hundreds of hours painting a great artwork and then deciding that everyone should be able to give a copy to everyone they know, under the simple condition that they inform those people that they have this right as well.
And then comes along a company and sells my artwork for money, without informing their customers that they can pass it on for free. That’s, plain and simple, a criminal operation.deleted by creator
No one. What people give a shit about is the license that is supposed to keep the code open, which is being removed for profit without consequence.
In our current legal system, copyright is the basis for me to be able to set requirements on how my code can be shared. I do not care that I own it, I just care that it is shared under the conditions I set.
Without being able set these conditions, I would not open up my code.
If nobody owns the code, then then nobody can enforce the terms of the license it was released under, and free software under the FSF definition becomes impossible. All you have is public domain.
For example, a company could take the Linux kernel, modify it and distribute it with their gadgets. And they could simply not release the modifications they’ve made, as is required by the GNU Public License. But nobody would be able to do anything about it. Currently, copyright laws allow the people who wrote the Linux kernel to sue the company for breaking the license and violating the authors’ copyrights
Noteworthy: They crawled only the default branch HEAD and inlined all source content.
- The file contents are included inline. The decoded UTF-8 source text is embedded directly in the dataset, so it is fully self-contained — you can start training the moment the download finishes.
- It reflects the state of GitHub in August 2025. The corpus is a direct crawl of GitHub repositories at their default-branch HEAD, capturing roughly two additional years of open-source code compared to The Stack v2.
they can do whatever they want to do with my code.
as long as it’s complaint to AGPL v3 :D
(doubt ai companies care about that)
Scraping without consent steals developer personalIP.
Is there a form to ask to be included in the next stack? they seem to have missed me this time













