Responsible use of llms can’t be achieved. Can you verify that all of your training data are square with its creators? If not, how is that responsible?
Can you verify that it didn’t reproduce any code that’s proprietary or has more restrictive licenses than your own?
That’s the main issue for FOSS.
The complement/ reverse is true for companies that produce proprietary software (but they won’t care, sure.)
It’s possible, but they’re not doing it. They chose to be irresponsible
You can’t verify that with humans either. Someone who has seen licensed code, whether proprietary or GPL, will accidentally reproduce snippets of it to accomplish similar tasks in the future.
The problem isnt the fact snippets exist in a codebase. The problem is, the entire fucking website is scraped clean of every codebase that has ever existed without paying a dime to their owner, notifying of their usage, and being transparent of their datasets.
How can it be verifiably responsible if their activity is handwaved under the guise of “business secret”? Its literally people using other peoples work without fair compensation, which means stealing, which means someone has to pay.
The problem isnt the fact snippets exist in a codebase.
No, that IS the problem legally, but it goes slightly beyond snippets, it’s also about having similar overall design.
The problem is, the entire fucking website is scraped clean of every codebase that has ever existed without paying a dime to their owner, notifying of their usage, and being transparent of their datasets.
That is also how humans work. Any code you write is going to be subconsciously influenced by code you’ve seen before, especially code you yourself have written in other projects that you may not legally own. Much like an LLM, the human brain is a black box in that you don’t always know where an idea or something comes from.
This is why clean-room designs are often necessary. You need to replicate something a GPL project does, you should only use devs that have never seen that GPL project’s code. Same for replicating proprietary projects, you don’t use people that have seen that project’s’ code. So again, either the whole issue is stupid and we can use LLMs too, or it’s a real issue and everything should be clean-roomed to avoid any chance of an accidental licensing issue.
Proprietary projects and clean room implementation is often done without having the source code first though, and a clean room implementation assumes its actually clean. meanwhile, most of llms nowadays cant be verifiably do clean room because the source code is probably there too.
When Linus does clean room to make Linux, does he have access to the source code, or does he need to figure out things by himself and only match the interface later on?
I’m sure both of us aren’t good enough to solve this billion dollar problem, but saying its not a problem to steal someone’s work, disregarding their license (even MIT requires attributions!), and releasing their mangled copy as a new product, is dystopian.
In a nutshell, it looks like AI disclosures are encouraged, but not required, and AI usage is encouraged to be human reviewed, but not required. They have also stated they will not allow the use of online AI services for security reasons (but how this will be enforced I’m not sure, since they are relying on the judgement of contributors)
No online services? So you can use AI to generate code, but only garbage local AIs (assuming you don’t have a ton of RAM)?
This seems like the weakest possible decision they could’ve made.
It’s open source software. I’m not saying security isn’t a concern, but this is just stupid.
If you’re not going to ban AI then you should at least take advantage of models that are more reliable and produce higher quality output.
It looks like their point is that they don’t want Debian’s codebase to be used to train corporate AI models, and almost all the proposals seem to agree on that at the very least. I feel like a required AI disclosure would have been better, but what do I know, I’m not a Debian contributor
I am moving to Piefed at scallisto@retrofed.com (https://retrofed.com/u/scallisto)
I mean the best local models are only a few months behind the the best proprietary cloud models. Just look at Qwen 3.8 Flash or 27B. Imo local models are more than sufficient if you’re going to use AI for programming. Who cares if the small model can’t oneshot the whole patch, you shouldn’t be submitting raw LLM code to public repos anyway. Imo the only acceptable use of LLMs for coding is as a way of rapidly prototyping ideas that you will later mostly/entirely rewrite by hand, or as an extra static analysis tool for finding potential security holes.
And as someone who frankly doesn’t give a shit about intellectual property over code, my main ethical issue with LLMs is the monstrous resource consumption of hyperscale datacenters, so local models are strongly preferable. Also for privacy reasons.
They’re getting better fast
https://www.xda-developers.com/qwen-3-8-27b-reverse-engineering-job-frontier-model/
Still unreliable for tasks which aren’t easy to validate in code, though
The only question I have is why. Genuinely.
Debian exists. It has existed for decades. It works. “Oh but AI makes it faster” so fucking what? We didn’t need “faster” for years whilst it worked fine, it’s not a product on a deadline.
No community distro has the capacity to fork thousands of packages and maintain them. That’s what a true no Ai policy would require. I also have to add that the Linux kernel itself would need to be forked and then maintained. Just the latter point alone is enough to make all this moot.
Completely missing the point.
No this was for Debian only.
Yikes, rest in peace Deb users. I just hope it never happens to my distro.
I ain’t shook up about it. There’s not really a surefire way to detect tool-assisted code gen anyway, so IMO the acceptance criteria should be the same as it’s always been, tool-assisted or otherwise. Which is ultimately the path they chose to take.
Obvious slop should be immediate permanently banworthy, sloppers can just keep burning new accounts (as long as they have access to new IP addresses) while real developers deserving of praise and reputation thrive.
sloppers can just keep burning new accounts (as long as they have access to new IP addresses) while real developers deserving of praise and reputation thrive.
Becoming a Debian developer requires you to meet an existing Debian developer in person and have your public key signed by them. It’s not possible to keep burning new accounts unless you go and meet a different Debian developer each time and there’s a limited number of them in each region and they usually meet together, so more than one person will see your face.
Yep. I’m not familiar with Debian’s strategy specifically, but generally, I think any new contributor to a project should be subject to heightened scrutiny. My policy is that new contributors should start small and develop a rapport with the maintainers before submitting more ambitious (and for the maintainers, more costly to review) large and/or critical path PRs. It was a good policy before LLMs and I think it remains a pretty robust method of weeding out irresponsible devs without wasting a ton of maintainer time. There are simply more slop PRs to reject sight unseen these days, which is admittedly very annoying, but the process is much the same as it’s always been.
I’ll admit I don’t maintain any projects anywhere near the popularity or volume of the Debian project, so I’m not really sure what the view is from their vantage point.
I still think that’s not good enough, that treating them fairly is a stupid waste of time and resources and unfair to everyone else.
Just make the rule “any slop” and give the idiots a checkbox so they can ban themselves for reasons which will never be revealed to them (sloppers don’t read documents, it’ll take them a while to figure out). Also start banning people when evidence surfaces of them admitting to slopping.
I don’t like this concept that a slopper can potentially produce decent code, the data shows this simply isn’t true: sloppers produce vast amounts more and worse bugs and vulnerabilities. It’s better for the health of the project to ban it in every scenario.
How is treating one contributor fairly unfair to another contributor? If you want to add a “check this box to get your PR dumpster’d” checkbox I guess go nuts, but I’m unconvinced that’s a good long-term solution. I find it easier to ask “Do I know this contributor, or did they follow the new contributor guidelines and submit a small, single-issue PR?”, and if the answer is “no” then the PR gets ignored or, if I’m feeling gregarious and have the time, rejected with change requests. It’s a pretty easy rubric.
Humans have been perfectly capable of generating huge volumes of trash code, and code that looks good at first glance but has tricky bugs or vulnerabilities, since long before LLMs were a thing. The only real change now is the pace at which shitty code can be ripped out. IMO the solution is just: don’t accept more code than you can review and test. If that means rejecting 10x or 100x more LoC than you did five years ago, then… ok. It is more busy work, and it is annoying. But I don’t think trusting contributors to self-declare LLM use is an answer to the problem. There are better ways of rate-limiting eager beavers, regardless of what tools they use.
How is treating one contributor fairly unfair to another contributor?
Because sloppers do not actually do the work and in the vast majority of cases are not even capable of doing so. A slopper can produce 10 worthless products in the time a person can produce 1 good product.
Do you use arch btw?
At first I was afraid that they would allow for an irresponsible use. But no, they explicitly say “responsible”, so that danger is no more. I’m very much relieved.
Also they explicitly mention that humans will remain accountable. Not Nature, or Fate, or the gods; mark that. Good thinking there!
Problems solved.
for fuck sakes; stop pretending the csam generating machine is useful.
Sanity won imo. We can argue all day about what responsible Ai use is, but distro projects without allowing AI use at all and are serious about it are dead in the water. What are they supposed to d do? Disconnect with upstream?
There are more kernel forks every day, fuck the sloppers. From where I’m sitting the slop distros which allow or even rely on AI are sick with a cancer and doomed.
None of the proposals would have banned AI usage in upstream projects:
It’s not like the Debian policy teams have the power do that to remote dev teams anyway tho. At best they can discard an upstream done with AI.
“wah, they dead in the water if they ban the slop theft machine!”
I don’t fundamentally disagree with the sentiment, but local models are mostly benign as long as you actually understand the code you’re submitting and rewrite it to remove the slop portions. Plus prohibiting it entirely just means people will lie about using LLMs. I do think Debian should have mandated people disclose whether/how AI was used for any given commit tho
What about prohibiting people from looking at the leaked windows source code? It just encourages people to lie about it.
I don’t care if people look at the leaked Windows source code, frankly they should lie about doing so to avoid being sued. I don’t understand what your point is supposed to be.
Should wine ban it though, it just encourages people to lie 3:
Nooo, you see chosing utilitarian position is a sin.
Yep. There is a middle ground with generative AI and they’ve struck the right balance.
What were they supposed to d do? Disconnect with upstream?
Use non-slopified alternatives. There are lists such as Open Slopware that document where to get alternatives.
Or just fork the last non-slopified version and work from there. For already-mature software, should not make that much difference in the short term for a distro that seeks out stable long-term behaviour (things would be vastly different if this was Arch).
They basically pass the buck to the individual developer without taking any responsibility themselves.
Debian acknowledges that the legal status of material produced by generative AI systems remains the subject of ongoing discussion in many jurisdictions, including questions relating to copyright, authorship, licensing, and potential reproduction of training material.
The responsibility for every contribution rests with the contributor who submits it, who remains accountable for its technical quality, legal acceptability, and suitability for inclusion in Debian.
“It may be illegal or against FOSS, but that’s up to you to decide, good luck I guess”
This email from 2016 by jwz springs to mind.
I guess you want Debian to be the kind of operation that uses the work of others while blatantly and explicitly ignoring the wishes of the person who did the actual creative work.
I am increasingly of the opinion that all software developers and adjacent people are fucking scum unless proven otherwise.
Oh no, the people creating free shit for you to use that’s not monetized in any way, want to reduce their workloads. Scum!
You mean the work that that nobody is forcing them to do? The one they do by their own choice while not asking consent and respecting the wishes of others? Yep those devs sound pretty fucking scummy to me.
boooo
Dammit. Need to find a new OS.
If you want to avoid ai assisted code your only chance is to build a new OS with a new kernel yourself.
True. I’m really not happy with the avalanche of dogshit produced by “AI” though. Maybe I need to find a new hobby.
Not necessarily. No one says you can’t just use the latest commit to the Linux kernel (or any other software) before it became slopified.
i’m actually a little surprised, given their history about being so hardcore about dfsg compliance.
There is no such thing as “responsible use of GenAI”. That said, I understand the decision in the face of even Linus Torvalds allowing AI-generated code into the kernel. If Debian would have banned AI code entirely they would have had to fork every single thing they include in their distro, wipe the AI code from it, and then work to develop it further themselves. It would be too gargantuan a task.
None of the proposals would have banned AI usage in upstream projects, but banning AI usage in the context of the Debian project was on the table
There is. Local models trained only on clearly copy left licensed data.
But this isn’t that. This is very irresponsible
As long as they don’t allow AI bots to submit changes, this is probably a realistic decision. Assuming that code generated by AI is never going to be copyrighted by the AI companies. In light of this uncertainty I don’t understand why they don’t require AI code to be flagged as such. That bit seems like a really bad idea, legally speaking.
Good thing I don’t use debian.
















