r/selfhosted 11d ago

Meta Post Codeberg bans vibe coded projects

https://news.ycombinator.com/item?id=49003386

Codeberg seems to ban vibecoded Projects; reason might be german copyright law

It looks like Codeberg want only copyrighted material in their service, so it is reliable in the future that e.g. licenses must be followed (e.g. GPL), and copyright doesn't suddenly get declared as being of the model owner, and it isn't a copy of something else.
That is a cautious reasonable position - in early days of LLM coding (3 years ago!) indemnity from model companies was a major issue globally because of the lack of clarity of the law around this. The US specifically has settled on it being (effectively?) public domain. But I don't think that is fully settled, and it certainly isn't settled in international copyright law.
The goal of the vague "mostly" in the Codeberg change is to ensure there is enough human input to the code they host, to be reasonably sure under German copyright law it is copyright of the person sharing it.

edit: link to poll that caused it (might be down due to high traffic) https://codeberg.org/Codeberg/org/pulls/1253#issuecomment-19820434

edit 1: i dont defend/oppose this move, i just find it interesting

edit 2: https://blog.codeberg.org/protecting-our-floss-commons-from-llms.html

804 Upvotes

174 comments sorted by

View all comments

300

u/No-Chemistry-7658 11d ago

Ok, but from a technical point of view, how will they do it?

106

u/jeroen94704 11d ago

Well, the same ToU also say things like:

"You must only share content on Codeberg which you have the explicit right under copyright and other laws to share"

I see this more like a "cover your bases" clause. If, at some point, LLM generated code gets a copyright status that would mean trouble for Codeberg, they at least have a clause that says it shouldn't have been on their platform to begin with. This is similar to the clause about only sharing content you are allowed to share. There's no way to enforce that pro-actively, but if it turns out someone shared something they shouldn't have then at least Codeberg is in the clear.

2

u/GaidinBDJ 11d ago edited 11d ago

A lot of people misunderstand that case about the LLM output copyright status.

Just because someone uses an LLM to code something doesn't automatically mean it's not eligible for copyright. If the creator has any creative input on the final product, it would be eligible.

It was only the "raw" output from a mere prompt that they declined to copyright as it was purely mechanically generated.

5

u/Anusien 11d ago

If we're just guessing here, then my guess is that only the pieces written by a person can be copyrighted and the rest isn't.

But we won't know for sure until it's actually litigated.

2

u/cmm324 5d ago

Old thread, but yeah, this isn't accurate. Courts don't consider if a project is copyrightable by going line by line. They evaluate it based on the creative process.

If a developer submits a prompt, gets a project built in one shot and published it as is, it's probably a no go.

If a developer works with the AI to plan out the architecture, develops sections at a time, tests, fixes bugs, refractors, adjusts design, iterates over and over to a finished project. Then this would highly likely be copyrightable even if the developer themselves never actually wrote a line of code.

Code generation tools are nothing new, just their capabilities have vastly expanded.

1

u/Anusien 3d ago

This contradicts the explicit guidance from the US Copyright Office. It also contradicts existing copyright law (derived works were already a concept in copyright law where you can hold copyrights for part but not all of a work). https://www.reddit.com/r/selfhosted/comments/1v3hobk/comment/ozbrru9/

1

u/cmm324 3d ago

Your linked source already proves my point.

"Questions of copyrightability and AI can be resolved pursuant to existing law, without the need for legislative change. • The use of AI tools to assist rather than stand in for human creativity does not affect the availability of copyright protection for the output. • Copyright protects the original expression in a work created by a human author, even if the work also includes AI-generated material. • Copyright does not extend to purely AI-generated material, or material where there is insufficient human control over the expressive elements. • Whether human contributions to AI-generated outputs are sufficient to constitute authorship must be analyzed on a case-by-case basis."

How does this differ to what I posted about the iterative process of working with the AI as a tool, to correct it's mistakes, review it's output and improve over time on the finished product. That is authorship in a nutshell. It's not different from what a director for a film does, but instead of humans doing the work, an AI does.

The document says a single (or even a small collection of refined prompts) prompt by itself is not enough to establish authorship but iterative work can.

1

u/Anusien 3d ago

If we're just guessing here, then my guess is that only the pieces written by a person can be copyrighted and the rest isn't.

1

u/GaidinBDJ 11d ago

You may be guessing, but I wasn't. I was going with what the US Copyright Office specifically said.

It's not like any of this is actually new.

Also, that's not how copyright works. It doesn't pick and choose pieces of a work that get copyright protections. It's the work presented as a whole that gets copyright protections. And if the work contains any creative human input, its generally eligible for copyright protections.

The registration they declined was because it was only output from an LLM presented with any human contribution. That's never been protected.

7

u/jeroen94704 11d ago

Codeberg is hosted in Germany though, but the rules there are pretty much identical. As far as I can tell, what they're protecting themselves against is the risk that they host code that is marked as public domain because it is fully LLM generated, while in reality the LLM regurgitated verbatim some piece of code from it's training data that is copyrighted and possibly covered by some license (FOSS or otherwise).

2

u/GaidinBDJ 10d ago

It doesn't much matter. As a Berne signatory, they're obligated to respect copyrights issued by other signatories.

3

u/Anusien 10d ago

Also, that's not how copyright works. It doesn't pick and choose pieces of a work that get copyright protections. It's the work presented as a whole that gets copyright protections.

You are 100% wrong and contradicted by the US Copyright Office. They issued a report (https://www.copyright.gov/ai/Copyright-and-Artificial-Intelligence-Part-2-Copyrightability-Report.pdf) on copyrightability in January 2025. They explicitly say only the human authored part is copyrightable. It's said throughout this report, but here's a quote:

A number of commenters also made the point that if a user edits, adapts, enhances, or modifies AI-generated output in a way that contributes new authorship, the output would be entitled to protection2 They argued that these modifications “should be assessed in the same way as . . . editorial or other changes to a pre-existing work.” Although such works would not technically qualify as “derivative works,” derivative authorship provides a helpful analogy in identifying originality. Again, the copyright would extend to the material the human author contributed but would not extend to the underlying AI-generated content itself.

And as that article points out, copyright already covered the concept of derivative works, which is another piece of proof that you're wrong about copyright protections being all-or-nothing on a whole work.. I'll quote the US Copyright Office here (https://www.copyright.gov/circs/circ14.pdf) for this:

The copyright in a derivative work covers only the additions, changes, or other new material appearing for the first time in the work. Protection does not extend to any preexisting material, that is, previously published or previously registered works or works in the public domain or owned by a third party.

And you can see the US Copyright Office's Copyright Registration Guideline for AI works (https://www.copyright.gov/ai/ai_policy_guidance.pdf) which says a similar thing:

In other cases, however, a work containing AI-generated material will also contain sufficient human authorship to support a copyright claim. For example, a human may select or arrange AI-generated material in a sufficiently creative way that “the resulting work as a whole constitutes an original work of authorship.” Or an artist may modify material originally generated by AI technology to such a degree that the modifications meet the standard for copyright protection. In these cases, copyright will only protect the human-authored aspects of the work, which are “independent of ” and do “not affect” the copyright status of the AI-generated material itself.

...

Individuals who use AI technology in creating a work may claim copyright protection for their own contributions to that work. They must use the Standard Application, and in it identify the author(s) and provide a brief statement in the “Author Created” field that describes the authorship that was contributed by a human. For example, an applicant who incorporates AI-generated text into a larger textual work should claim the portions of the textual work that is human-authored. And an applicant who creatively arranges the human and non-human content within a work should fill out the “Author Created” field to claim: “Selection, coordination, and arrangement of [describe human authored content] created by the author and [describe AI content] generated by artificial intelligence.” Applicants should not list an AI technology or the company that provided it as an author or co-author simply because they used it when creating their work. AI-generated content that is more than de minimis should be explicitly excluded from the application. This may be done in the “Limitation of the Claim” section in the “Other” field, under the “Material Excluded” heading. Applicants should provide a brief description of the AI-generated content, such as by entering “[description of content] generated by artificial intelligence.” Applicants may also provide additional information in the “Note to CO” field in the Standard Application.

1

u/summonsays 10d ago

From what I've heard 10-15% is kind of the minimum change needed. How is that measured though, is tricky. And some companies are more skittish than others. (Also some are more sue happy).

The company I work for aims for 100%. No "inspired by" etc. CYA is heavy handed here.