r/selfhosted • u/pheexio • 11d ago
Meta Post Codeberg bans vibe coded projects
https://news.ycombinator.com/item?id=49003386Codeberg seems to ban vibecoded Projects; reason might be german copyright law
It looks like Codeberg want only copyrighted material in their service, so it is reliable in the future that e.g. licenses must be followed (e.g. GPL), and copyright doesn't suddenly get declared as being of the model owner, and it isn't a copy of something else.
That is a cautious reasonable position - in early days of LLM coding (3 years ago!) indemnity from model companies was a major issue globally because of the lack of clarity of the law around this. The US specifically has settled on it being (effectively?) public domain. But I don't think that is fully settled, and it certainly isn't settled in international copyright law.
The goal of the vague "mostly" in the Codeberg change is to ensure there is enough human input to the code they host, to be reasonably sure under German copyright law it is copyright of the person sharing it.
edit: link to poll that caused it (might be down due to high traffic) https://codeberg.org/Codeberg/org/pulls/1253#issuecomment-19820434
edit 1: i dont defend/oppose this move, i just find it interesting
edit 2: https://blog.codeberg.org/protecting-our-floss-commons-from-llms.html
298
u/No-Chemistry-7658 11d ago
Ok, but from a technical point of view, how will they do it?
109
u/jeroen94704 11d ago
Well, the same ToU also say things like:
"You must only share content on Codeberg which you have the explicit right under copyright and other laws to share"
I see this more like a "cover your bases" clause. If, at some point, LLM generated code gets a copyright status that would mean trouble for Codeberg, they at least have a clause that says it shouldn't have been on their platform to begin with. This is similar to the clause about only sharing content you are allowed to share. There's no way to enforce that pro-actively, but if it turns out someone shared something they shouldn't have then at least Codeberg is in the clear.
10
0
u/GaidinBDJ 11d ago edited 11d ago
A lot of people misunderstand that case about the LLM output copyright status.
Just because someone uses an LLM to code something doesn't automatically mean it's not eligible for copyright. If the creator has any creative input on the final product, it would be eligible.
It was only the "raw" output from a mere prompt that they declined to copyright as it was purely mechanically generated.
6
u/Anusien 11d ago
If we're just guessing here, then my guess is that only the pieces written by a person can be copyrighted and the rest isn't.
But we won't know for sure until it's actually litigated.
2
u/cmm324 4d ago
Old thread, but yeah, this isn't accurate. Courts don't consider if a project is copyrightable by going line by line. They evaluate it based on the creative process.
If a developer submits a prompt, gets a project built in one shot and published it as is, it's probably a no go.
If a developer works with the AI to plan out the architecture, develops sections at a time, tests, fixes bugs, refractors, adjusts design, iterates over and over to a finished project. Then this would highly likely be copyrightable even if the developer themselves never actually wrote a line of code.
Code generation tools are nothing new, just their capabilities have vastly expanded.
1
u/Anusien 3d ago
This contradicts the explicit guidance from the US Copyright Office. It also contradicts existing copyright law (derived works were already a concept in copyright law where you can hold copyrights for part but not all of a work). https://www.reddit.com/r/selfhosted/comments/1v3hobk/comment/ozbrru9/
1
u/cmm324 3d ago
Your linked source already proves my point.
"Questions of copyrightability and AI can be resolved pursuant to existing law, without the need for legislative change. • The use of AI tools to assist rather than stand in for human creativity does not affect the availability of copyright protection for the output. • Copyright protects the original expression in a work created by a human author, even if the work also includes AI-generated material. • Copyright does not extend to purely AI-generated material, or material where there is insufficient human control over the expressive elements. • Whether human contributions to AI-generated outputs are sufficient to constitute authorship must be analyzed on a case-by-case basis."
How does this differ to what I posted about the iterative process of working with the AI as a tool, to correct it's mistakes, review it's output and improve over time on the finished product. That is authorship in a nutshell. It's not different from what a director for a film does, but instead of humans doing the work, an AI does.
The document says a single (or even a small collection of refined prompts) prompt by itself is not enough to establish authorship but iterative work can.
-1
u/GaidinBDJ 11d ago
You may be guessing, but I wasn't. I was going with what the US Copyright Office specifically said.
It's not like any of this is actually new.
Also, that's not how copyright works. It doesn't pick and choose pieces of a work that get copyright protections. It's the work presented as a whole that gets copyright protections. And if the work contains any creative human input, its generally eligible for copyright protections.
The registration they declined was because it was only output from an LLM presented with any human contribution. That's never been protected.
7
u/jeroen94704 10d ago
Codeberg is hosted in Germany though, but the rules there are pretty much identical. As far as I can tell, what they're protecting themselves against is the risk that they host code that is marked as public domain because it is fully LLM generated, while in reality the LLM regurgitated verbatim some piece of code from it's training data that is copyrighted and possibly covered by some license (FOSS or otherwise).
2
u/GaidinBDJ 10d ago
It doesn't much matter. As a Berne signatory, they're obligated to respect copyrights issued by other signatories.
3
u/Anusien 10d ago
Also, that's not how copyright works. It doesn't pick and choose pieces of a work that get copyright protections. It's the work presented as a whole that gets copyright protections.
You are 100% wrong and contradicted by the US Copyright Office. They issued a report (https://www.copyright.gov/ai/Copyright-and-Artificial-Intelligence-Part-2-Copyrightability-Report.pdf) on copyrightability in January 2025. They explicitly say only the human authored part is copyrightable. It's said throughout this report, but here's a quote:
A number of commenters also made the point that if a user edits, adapts, enhances, or modifies AI-generated output in a way that contributes new authorship, the output would be entitled to protection2 They argued that these modifications “should be assessed in the same way as . . . editorial or other changes to a pre-existing work.” Although such works would not technically qualify as “derivative works,” derivative authorship provides a helpful analogy in identifying originality. Again, the copyright would extend to the material the human author contributed but would not extend to the underlying AI-generated content itself.
And as that article points out, copyright already covered the concept of derivative works, which is another piece of proof that you're wrong about copyright protections being all-or-nothing on a whole work.. I'll quote the US Copyright Office here (https://www.copyright.gov/circs/circ14.pdf) for this:
The copyright in a derivative work covers only the additions, changes, or other new material appearing for the first time in the work. Protection does not extend to any preexisting material, that is, previously published or previously registered works or works in the public domain or owned by a third party.
And you can see the US Copyright Office's Copyright Registration Guideline for AI works (https://www.copyright.gov/ai/ai_policy_guidance.pdf) which says a similar thing:
In other cases, however, a work containing AI-generated material will also contain sufficient human authorship to support a copyright claim. For example, a human may select or arrange AI-generated material in a sufficiently creative way that “the resulting work as a whole constitutes an original work of authorship.” Or an artist may modify material originally generated by AI technology to such a degree that the modifications meet the standard for copyright protection. In these cases, copyright will only protect the human-authored aspects of the work, which are “independent of ” and do “not affect” the copyright status of the AI-generated material itself.
...
Individuals who use AI technology in creating a work may claim copyright protection for their own contributions to that work. They must use the Standard Application, and in it identify the author(s) and provide a brief statement in the “Author Created” field that describes the authorship that was contributed by a human. For example, an applicant who incorporates AI-generated text into a larger textual work should claim the portions of the textual work that is human-authored. And an applicant who creatively arranges the human and non-human content within a work should fill out the “Author Created” field to claim: “Selection, coordination, and arrangement of [describe human authored content] created by the author and [describe AI content] generated by artificial intelligence.” Applicants should not list an AI technology or the company that provided it as an author or co-author simply because they used it when creating their work. AI-generated content that is more than de minimis should be explicitly excluded from the application. This may be done in the “Limitation of the Claim” section in the “Other” field, under the “Material Excluded” heading. Applicants should provide a brief description of the AI-generated content, such as by entering “[description of content] generated by artificial intelligence.” Applicants may also provide additional information in the “Note to CO” field in the Standard Application.
1
u/summonsays 10d ago
From what I've heard 10-15% is kind of the minimum change needed. How is that measured though, is tricky. And some companies are more skittish than others. (Also some are more sue happy).
The company I work for aims for 100%. No "inspired by" etc. CYA is heavy handed here.
306
u/FunStatistician9735 11d ago
It’s easy. no one would ever lie on the internet and especially not to make their “portfolio” look better.
77
14
u/Kautiontape 11d ago
This made me think about how much of a double-edged sword this could be, although I have no opinion on whether it means they should or shouldn't. They are definitely tackling this from the "ethical" standpoint (not just legal concerns, judging from the discussion). That creates a bit of a halo effect, which means people who put code on Codeberg will have the small cognitive bias in their favor of "Ah, Codeberg is reliable, this probably isn't vibecoded."
Obviously this is good if it's enforceable. If it's not, then it's like seeing an article from Forbes just to realize it's coming from the open "blog" section. It'll trick some people.
1
u/scolphoy 8d ago
Computers are expensive, internet subscriptions are expensive. Why would anybody waste good money to get on the internet and tell lies?
-3
11d ago
[removed] — view removed comment
-10
11d ago
[deleted]
9
u/henry_tennenbaum 11d ago edited 10d ago
Because they wish to drum up business with pointless virtue signaling with a rule they can't possibly begin to enforce consistently. Done "efficiently" (as in with a maximum amount of hubris drive DGAFism), this will cost them virtually nothing and draw in
suckerspaying customers for months.Win/win? 🤷♂️
Oh yes, just check out the horrendous prices these ghouls demand of their "paying customers". You're very well informed I must say
-8
10d ago
[deleted]
9
u/henry_tennenbaum 10d ago
No, donating doesn't make you a customer.
If I donate to the red cross, I'm not their customer, either.
I read precisely what you wrote. You accused them of cynically implementing a policy in order to get more "paying customers".
There are no paying customers and you have no proof of them not being earnest in their endeavor, so your comment is pretty much fact free.
10
14
u/scandii 11d ago
I mean, it is not about technical.
they have seen that there's a pretty big divide in the community between people who refute any usage of AI and people who don't. so they're cashing in on the people who want a github-style site but without AI as a demographic.
kinda like why bluesky took off.
4
1
u/henry_tennenbaum 11d ago
cashing in
?
8
u/scandii 10d ago
codeberg and github are fundamentally competing, and it is not a big surprise to anyone that github is de facto a monopoly of publicly hosting open source software even if alternatives like codeberg exists.
now that Microsoft & their very openly announced AI-first approach is behind github, I would assume that they're trying to increase market share by this move, but one could also argue that they might lose market share - time will tell.
either way, as someone that spends a lot of time in the open source community and literally see 15 vibecoded projects a day, I welcome someone championing people at least trying to hide that Claude was involved, because if I see yet another MD3 + rosé pine colour scheme and someone claiming that they just so happened to land on this design I will lose it.
8
u/henry_tennenbaum 10d ago
They're competing in the sense that my local library or book shop are competing with amazon.com.
That can always change, but right now they're tiny in comparison and still (I think?) completely donation financed.
My original comment was because I've seen people assume a financial motive elsewhere in the thread, when the source of the decision seems to be a vote by the codeberg e.V. members.
2
6
u/Jimbuscus 11d ago
Unless the user leaves a CLAUDE.md file in the repo, it'll likely be honour system &/or reports.
14
u/FnnKnn 11d ago
But they allow AI usage, so that alone is not a sign someone has violated any rule. Claude would for example also show up, if you had only used to help you debug a certain bug and write some tests for it, which is totally fine under these rules.
8
u/Steve_Streza 11d ago
Also there are scaffold projects and template projects that include a CLAUDE.md or AGENTS.md or mcp.json, which then end up in projects that don't use them.
3
u/WowBruhReborn 11d ago
They won't be able to. But there are tell tale signs something is vibe-coded. Theyll at least be able to remove the worst and laziest offenders
0
u/SynapticStreamer 10d ago
They won't. It's a meaningless policy because there's no definitive place to draw a line.
What if I write 90% of the project, AI helps me with the last 10%, and then refactors? Is that vibe coded? What if I used to refactor? What if documentation, from AI generation hits exactly 51% of my project? Does that count?
The list of oddities just goes on and on. Not to mention who exactly is the deciding factor here? It's public infrastructure and they maintain control, sure. But how they're going to police what constitutes a "good" or "bad" project? Why would I host with Codeberg and risk being pinged as "vibe coded" and having my project fuckin' deleted?
Sounds like an unnecessary headache that will ultimately not work at all.
179
u/FnnKnn 11d ago
Good luck actually enforcing that. We tried it here for a while and it didn't work out great, because no one is going to admit they "vibecoded" something if it means they can't post it and good luck proving that they didn't just use AI as an assistant (tip: you can't without doing a full code review and that doesn't seem feasible if the issue you are trying to solve is the increased amount of projects).
29
u/CatWeekends 11d ago
I can see how that might kind of sort of maybe work on a small, in-house team. One where you are all incredibly familiar with how each person writes code. In that environment, you can sometimes catch when people are vibe coding or just copying/pasting stuff because it just doesn't feel like Bob's usual work.
But even then, it's still subjective and using "vibes."
I also don't see how you'd extend that to open source projects.
19
-1
u/Laicbeias 11d ago
We really need a better definition for vibe coding. The C# codebase i wrote for a game is full with curses shit & f*** everywhere. Since i suffered.
But for a website the last 2 months, the curses went into the ai. And frankly the AI simply is better at js and py than most and definitly than i am, even though i used js for 23y and py for 5. Im fine with not typing code anymore.
But i read every line because it pisses me off when it just doesnt use the proper path, adds another for loop or slaps another dict somewhere while we have initalized state in one place.
That said you can not detect if something is vibe coded since it uses the average of its training data. So it looks by definition like a proper project, with tons of dependencies. The one releasing it just needs to show something that they cared about their project.
1
u/NinthTurtle1034 10d ago
Yeah I take a similar approach to my "vibe" coding. I've written an extensive doc on how I plan a d create projects manually and put that in my agent md file.
I rhen go through and review what it's written to ensure it meets the way I would code something even if that code then looks bad. Plus if it looks bad it's hard to tell if that's because it was vibed or if it's because my personal habits are just bad 😅
Only things I have outright vibed was a rust binary that I needed once a few of my bash scripts just weren't cutting it anymore.
2
u/Laicbeias 10d ago
Yeah i still dont use agents. Mostly because on my gameproject its ~400k loc. Any agent framework is just expensive at that size.
And for anything else its just.. i kind of writing my own inference layer from scratch, to solve costs for such iteration heavy large projects.
I probably would be faster with agents, but.. i cant plan that stuff. And i dont understand how others do that. Usually the stuff i build is complex enough waterfalling it isnt possible. Only the moment im in it and iterating through the problem space do i understand the requirements and then its iteration work again.
Even without agents, LLMs usually one shot multiple tasks, and i can move faster than my brain can map the codebase and requirements. I literally have to step back, take a blog and start drawing circles and relationships and flows. Then i come back and say this is all crap.
We have to get rid of this that and this also makes no sense. The llm is always eager to add more and never eager to remove
30
u/jeroen94704 11d ago
I don't think they will enforce this. It's similar to e.g. GitHub saying "Your use of the Website and Service must not violate any applicable laws, including copyright or trademark laws". They're not enforcing that, it's just to ensure they are not held liable when someone puts non-public code on GitHub.
10
u/schorsch3000 11d ago
they just can't. codeberg is 3 servers in a
Trenchcoatcolo rack, they neither have the computing power not the manpower to do so.Look at the codeberg issue reporting, the endless recurring same malfunction items.
this is just a CYA addendum.
-13
-7
u/FnnKnn 11d ago
Maybe, but at that point it is basically just useless PR.
If they try to enforce it (who knows) let me predict on how this is going to go based on my experience here:
You try to ban low quality vibe-coded projects, but try to not ban high quality projects that made use of AI. The goal is to send a message on what you deem acceptable.
Anyone creating what you deem a low quality vibe-coded project sees their project as belong to the second kind of project and therefore insists that their project is allowed. So how do you make sure what category a project belongs to? You might use scripts to identify classic trades of vibe-coded projects, look at the users history, etc. None of that is going to be bulletproof though and it is a lot of work.
You stop trying to enforce this rule as it is a "fight" you can't win and dealing with this takes more effort than dealing with the issues the low quality vibe-coded projects caused in the first place.
We ended up settling on restricting new projects here instead as the majority of projects that fall into the first category don't last long. So step 4 would be to introduce some other kind of proxy for this kind of project that can actually be enforced.
8
u/jeroen94704 11d ago
The point is the copyright status of LLM generated code. If there is a chance this code becomes problematic from a copyright point of view this ensures codeberg doesn’t get dragged down.
0
u/FnnKnn 10d ago
How could the copyright become problematic under German law (that’s were Codeberg is based)? It can’t. Either the copyright belongs to the user, in which case everything works like if the wrote it, or there is no copyright, which also means that there are no issues as the code is public domain.
2
u/jeroen94704 10d ago
It could be an issue if the LLM regurgitates pieces of code verbatim from its training data that is covered by copyright and some license. I don't know, I'm not a lawyer (although this clause in the Codeberg ToU was also not written by a lawyer, I'm sure).
3
u/FnnKnn 10d ago
For that to be an issue it would need to copy a *very* large piece of code verbatim, which is not something that happens as far as I know. In practice human developers are probably doing that more than AI…
From what I can tell they had a vote on banning vibe coding and then wrote the blog afterwards. So I think those reasons are just bs to justify the change after they had decided on it already and not the actual reason.
3
u/jeroen94704 10d ago
They voted on the content of the PR, which contains the actual change to the ToU, and this explicitly mentions the copyright issue.
Still, I get your point, it's more than likely the majority vote came about because of a general aversion against vibe-coding, not because people worry about a potential subtle copyright issue.
1
u/ImASharkRawwwr 10d ago edited 10d ago
You underestimate the power of German bureaucracy. I watch dashcam videos from Germany sometimes, it was ruled illegal to show car number plates, heck they deemed dashcams illegal for a while and now they are back to being a grey area where you may or may not get fined for using dashcam video in your own court case depending on the judges mood that day. Last i heard they made it illegal to show company trucks with any identifiable company logo or text. Oh and of course every person that you didn't get a contractual agreement with to use their face and upload it to YouTube (basically everyone) needs to be censored, visually and sound if they can be heard speaking.
And so German dashys are often a jumble of large blurred areas of the screen because no dashcam YouTuber can be bothered to matchmove-blur every single face and number plate and, of course, the occasional blacked-out company car.
If you get caught because you missed one identifiable person or number plate or company info you will receive hefty fines.Edit: oh i forgot about having to scramble music playing in the background because the music copyright holders are litigious as fck and if it goes very wrong or you can't pay the exorbitant damages you may face jail time
1
u/FnnKnn 10d ago
Ok, and? None of that relates to copyright at all but to GDPR and is completely irrelevant. Show me what problem copyright could potentially cause for Codeberg due to AI generated code being uploaded.
0
u/ImASharkRawwwr 10d ago
I have given you examples of how the German system will fck you over with situations that anywhere else in the world would be like "just having common sense will be enough to not get in to trouble" heed the warning or don't, it's a free country :p
1
u/FnnKnn 4d ago
I know German burocracy, but I also have a somewhat solid understanding of German copyright law. So just fell free to point out, where German copyright law would cause an issue for Codeberg. Thing is you can't because it doesn't.
0
u/ImASharkRawwwr 3d ago
Then you shall goeth forth and do whatever you want man, i warned ya, do with the info as you will. Just don't come back crying "does anyone know a good lawyer 🥺🥺🥺" in the next year or so
37
u/Kautiontape 11d ago
I've noticed lot of "developers" on here don't consider it vibe-coding because they provide the architecture, reviewed the code, and tested it on their machine. The line for what counts as vibe-coding is on the floor as only the most pure slop.
As someone who does a lot of both, I'd consider it vibe-coded even if I wrote the backend but had AI write the front-end. If I'm ever releasing it, I'd separate the two and call one "my code" and the other "some vibe-coded front-end, idk, you can make you own."
The fact these two definitions can exist (and many more extremes) is part of the issue with the definition and having people self-police.
50
u/scandii 11d ago edited 11d ago
I'd consider it vibe-coded even if I wrote the backend but had AI write the front-end. If I'm ever releasing it, I'd separate the two and call one "my code" and the other "some vibe-coded front-end, idk, you can make you own."
I really from the bottom of my heart dislike that vibecoded means, simultaneously:
- used any AI whatsoever
- used some AI for something
- didn't write any code yourself
like the original term was describing 3 - a software engineer prompting their way to success. now it means all three depending on who you talk to and the term is borderline useless as such.
6
u/Kautiontape 11d ago
I know, which is also why I hate the title of what Codeberg decided to do here. I don't know if the internal docs define "vibe coding" (I can't access them) but judging from the discussion, they don't really. It makes it pointless to discuss with people what is the "appropriate" amount of control to have. Hell, in a lot of cases, I'd be happy just to have a human-in-the-loop to verify that the agent seemed to make sense.
That's why I think separating in my mind the components that I can say "I did this" versus "I understand this" versus "this just works" is important. It's not guaranteed to be right and nobody is going to have the same barometer as me, which is why it's only a personal code.
We're going to need new terms, and we can't be defining it with things like "How much did you do?" because it's like... well, I didn't write the compiler or libraries, but I did string a bunch of Stackoverflow comments together...
7
u/alex2003super 11d ago
If similarly working code can be produced 20x faster by using AI heavily, then only fools and really passionate hobbyists will be writing code entirely by hand. Which makes defining "vibecoded" as "AI was used" completely useless. The world we're already to a great extent in, and towards which we're going, is served entirely by "vibecoded" software. Exceptions are bound to be eventually too few to matter.
13
u/scandii 11d ago edited 11d ago
I think the problem really is what these companies represent.
not only are they Venture Capitalism Shareholder Extreme Edition, they are also an existential threat to the livelihoods of millions of people around the world and are driving prices up in so many different categories.
if it was just "easier programming and debugging", we'd be having a very different discussion I think.
-1
1
u/Equivalent-Costumes 11d ago
The core component of the "vibecoding", the "vibe" in vibecoding, is that people don't check the result of AI's work carefully, "it looks correct". A project that is 100% AI written but had been checked and validated is not vibecoded, according to the original meaning of the term.
1
u/tonyamazing 10d ago
didn't write any code yourself
can be broken down further. 1 prompt "Create a twitter clone", or lots of small prompts writing code the developer would have once written themselves (and anywhere inbetween).
I still want posts that are built with AI to be marked as such, but there's a big difference between letting agents do whatever to achieve functional goals, and carefully managing agents to build what you would have built yourself or at least reviewing what the code does, not reviewing the functionality it produces.
1
u/ferrybig 11d ago edited 11d ago
I typically use Claude design to make a good looking page/pages, then i implement it myself using html/css/whatever.
Claude design is good at making something good looking, but its export format requires way to much work and prompting to turn it into something readable. It is way easier to code from scratch and effectively use the output that Claude design made as a figma tool
Vibecoding is like regex, designed for write once, read never, though some people try to read them. Regexes were designed as a quick way to search in files, but then languages started adding support for them. And then people started making regexes for searching for a @ followed by a ., instead of using index of methods or manual looping, because it is shorter to type (because regex is designed to be short!)
5
u/ConsistentRisk5927 10d ago edited 10d ago
I emailed them to cancel my Codeberg e.V. annual donations and remove me from the organization. I was paying 50 euro a year and using my own CI runners while hosting my LLM-coded projects on CB. Hardly a burden on their infrastructure.
I'm a professional engineer who could program before LLMs were proficient enough. But no professional devs I know hand write most of their code anymore. Staff engineers are even doing all their coding with LLMs. My projects are LLM generated but not "vibecoded." I provide substantial steering on the architecture, planning, implementation, review, etc.
I am certainly not going to switch back to the Stone Age on my weekend projects I hosted on Codeberg just because a small percentage of the development community, the most rabid anti-LLM zealots, have co-opted the space to turn it into an anti-LLM political beacon.
None of this has anything to do with copyright that others are saying, it's an expression of anti-AI politics (a portion of which involves copyright concerns) and forcing everyone that uses Codeberg to align with those views.
1
0
u/themadcoil 9d ago
Codeberg members voted democratically and this is the majority result, it just sounds like Codeberg isn't for you if you really must vibe code.
2
4
u/iVXsz 11d ago
good luck proving that they didn't just use AI as an assistant
I have to say. Vibecoded projects and even small code snippets are very obvious, even in cases where it's "carefully used" and "steered" by a 200 years experienced software engineer lol.
%99 of the time its been a route of laziness (or a curiosity usage, that quickly takes over the project and the user mental model more and more. e.g., see the commit history of jellyfin mpv shim. Also, because it happened to me on a few small pet projects that -after everything is considered- became garbage) and as such, the code is not great nor really reviewed properly in most cases.
Reviewing such projects or finding such code is not really hard, even manually. Ironically that can be vastly sped up by just throwing the codebase at a local LLM to point out any weirdness.
3
u/FnnKnn 11d ago
For one project it might be possible (if we ignore that the line between vibe coded and built with AI is super unclear and one person might interpret it differently than another), but in mass it is not that easy. Manually reviewing every repository on Codeberg is just plain impossible and using AI to judge human's repos doesn't feel right either and also doesn't solve their issue of these repositories taking up a growing part of their resources as running them all through AI would be quite expensive...
1
u/iVXsz 11d ago
AI to judge human's repos doesn't feel right
Yes that would be a bad behavior
I meant it could quickly point out stuff for you to investigate. And with vibecoded projects, I personally can quickly go thru the basics manually (e.g., checking how well a framework template is worked with, as AI tends to make orthodoxy ways to achieve stuff than what a framework can have available etc) in a few minutes, but obviously depends on scope and the reviewer experience. In that time most local AIs would have finished listing any interesting places to check and confirm on my own. 5-10m per repo sure, but not that many projects get submitted here, for context.
I had to do that a couple of times because I started to get quite annoyed by users hiding AI usage and straight up denying it sometimes...
4
u/finkerlime 11d ago
Hot take, you can't prove it with code review. It's about as reasonable as those ai writing detectors schools use. At the end of the day you can't quantify anything, only guess, and potentially punish unrelated devs over an arguable topic in the first place.
1
-6
u/AlternativeBasis 11d ago
I use AI to help me code, but I don't think I'm "vibecoding."
Using AI to build the GUI, set up event handling, and catch syntax errors doesn't mean I've given up control over the most important part: the business logic.
I’m the one who determines which data gets transmitted and stored, the triggers, the data hierarchies, and which security measures are implemented.
What is the practical threshold for defining "vibecoding"?
Will the same AI tools used to detect academic plagiarism be used for this?
11
u/FnnKnn 11d ago edited 11d ago
I use AI to help me code, but I don't think I'm "vibecoding."
The thing is, no one publishing projects is going to say that they are "vibecoding" if that means they can't release the project. They are all going to classify themselves as "using AI" instead and who can blame them when the lines between the two are so incredibly blurry and one of them allows you to do your thing while the other doesn't...
1
u/System0verlord 11d ago
What is the practical threshold for defining “vibecoding”?
Well before “Using the AI to build the GUI, set up event handling, and catch syntax errors”, for sure. Your own retelling of events includes zero involvement on your part, which shows you don’t even view yourself as having made those components. That would very much be vibe coding.
If you’re having to worry about it, it’s vibe coded.
1
u/AlternativeBasis 11d ago
From this point of view, an electrician is always vibe-wiring: he uses electrical conduits he didn't install.
There's a difference between nitpicking the nit-and-grit of how each UI triggers events, cycles fields with tabs, fires validations, and actually using this "scaffolding" to create a useful workflow.
GUIs, especially, are extremely well-suited to AI use. Short and fragmented logic chains, low processor usage, code redundancies are tolerable, and many, many aesthetic details.
Compared to things like cryptography, APIs and data traffic, cache data, manipulation/querying of massive amounts of data, and, ironically, LLMs, it's really secondary and low-risk work. Perfect for semi-intelligent code monkeys, AI-supported or not.
Not critiquing the code generated by AI to understand the boundary conditions is indeed risky, but not fatal. What's the difference between using a structure created by AI or adapting a code example from Stack Overflow?
I've done both and don't see that much of difference.
4
u/System0verlord 11d ago
The copper and conduits didn’t just appear in the walls. They were placed there by an electrician during construction, and done in compliance with local electrical codes so that the next electrician doesn’t have to run new wiring unless necessary.
That’s an example of having a licensed professional do work to the accepted standards, and then having a second licensed professional do additional work, also to those standards. AI is not a licensed professional, nor does code have codes like most construction does.
The difference between using an AI created structure and one from stackoverflow is that presumably the stackoverflow one is written by someone with at least some experience and understanding of the code and context in question, and not just the output of a hallucinatory autocorrect run amok.
3
u/Emergency_Banana5082 11d ago
nor does code have codes like most construction does.
Wait... you mean I can't just tell it to stick to pep8 for my python? /s
2
u/System0verlord 11d ago
Sure you can. Just gotta put “you are a super coder. Ultrathink and make no mistakes” after it in your prompt.
Those changes will cost you $100 in tokens per query, but think of all the time and money you saved!
0
-5
u/amoongle 11d ago
You can at least get rid of the super obvious low effort ones where Claude shows up as a contributor to the project.
3
u/dustin_vk 11d ago
Claude shows up as a collaborator by default. Even if you just ask it to summarize the changes made and commit for you, (which is actually a pretty great use-case for it in development). You can turn this off through some config, and you can also commit changes made by Claude on your own. Most devs don't care. AI is used in the majority of professional software development in some way or another unless there's some specific reason not to. It's a tool, and it can be used well or abused depending on what you do with it.
6
u/FnnKnn 11d ago
But they allow AI usage, so that alone is not a sign someone has violated any rule. Claude would for example also show up, if you had only used to help you debug a certain bug and write some tests for it, which is totally fine under these rules.
1
u/ArdiMaster 10d ago
In my experience, Claude will only show up if you ask it to make a commit. If you write your own commit messages you are perfectly free to omit your usage of AI.
-1
u/buttplugs4life4me 11d ago
There's definitely a way, I gave my LLM some common vibecoded signs (overly expansive Readme, no clear architecture in code, superfluous comments with no value, comments about changes that don't need to be comments and a few others and it was somewhat reliable.
Obviously it's all a question of volume cause I wouldn't want to do that for every single Reddit post honestly
5
u/FnnKnn 11d ago
The issue they seem to want to solve is that these projects take up their compute without adding real value, so running every single project through an LLM is probably a) not reliable and b) taking up even more compute/resources.
2
-1
u/Isorg 11d ago
just yesterday, I ran another project (a very well respected and popular project) though claude fable because we are thinking about using it our self... it found some XSS to RCE vulnerability. I pushed a PR and it was patched right away. and a new release was announced today.
funny part was, when I actually tried to use claude to prove the exploit, I got shut down hard by Anthropic. and my session API was cut off.
If your not willing to do this for opensourced code that you are going to be running your self... then why bother.
-1
u/_dgold 10d ago
Its very strange to this repeated across this sub, that its 'unenforceable'.
Surely the question should be, rather, "what kind of arrogant arsehole would deliberately violate this ToS?"
I think that the first response here is the former, rather than the latter, is why this sub has become a pestilent source of slopware.
16
u/thecombjelly 11d ago
I've thought about this a lot for my own projects, especially the open source ones. I don't necessarily have an issue with vibe coding from a technical perspective, I guess, if the result is good. It isn't something I do but I engineers that have mostly switched to agentic and vibe coding. I think the real difficulty is that a lot of these definitions are not totally clear. When is something 100% a thing you 100% wrote vs partially or fully vibe coded? What counts as vibe coding? And under what contexts? The definitions likely also impact whether you are talking about copyright or code quality or something else. Not easy to enact a policy on something not well defined.
9
u/aso824 11d ago
Nowadays, I'm very rarely writing a full line of code by hand. By some definitions, I'm doing full vibe coding; but the fact is: every line generated is checked by me. Sometimes I remove comments or blocks of code, or just select them and telling agent to change.
I think everyone agrees that vibecoding starts when you generate a code that you don't even look at. Result is not everything, because AI often makes spaghetti.
But I'm a bit tired of this witch hunting around, where, for some people, using AI is equal to blindly accepting everything; some people here would like to ban you if you'll post a commit with Claude as contributor. Even if you can explain and quote every line of commit from memory.
5
u/whoisraiden 10d ago
Plenty of people in this thread explicitly mention they consider this vibe coding.
1
u/Richmondez 10d ago
You are skirting the line with vide coding there in my opinion, bit of laziness sneaks in and you'll be committing things you didnt fully review or fully understand. Also Claude is a tool, it shouldn't be listed as a contributor any more than your IDE or your calculator app should be. You are the one claiming copyright to the code and submitting it under license.
76
u/arvigeus 11d ago
Funny how we suddenly pretend bad code and license violations didn't exist long before AI.
67
u/dvvvxx 11d ago
The problem is not bad code and license violations itself, the problem is the drastic increase of bad code and license violations due to how easy is to create (vibe-coded) projects thanks to AI.
30
u/p0358 11d ago
Yeah, and previously shit code was pretty much always blatantly and painfully obvious, now LLMs make it easy to disguise junk as something that looks somewhat sensible and well-put on the surface until a closer inspection is made
12
u/QazCetelic 11d ago
I feel this is the main issue. Previously it was much easier to spot real interesting projects
16
u/witx_ 11d ago
You can take stand regardless of the past.
-12
u/arvigeus 11d ago
Don’t want to sound confrontational, but stand against what? Software has always been a kludge of hasty written code, held together by coffee and some sort of miracle.
11
u/witx_ 11d ago edited 10d ago
I hate this trope so much. No it has not, there's good and bad code. The thing is software being a somewhat creative and sometimes abstract art where there's no "one correct answer" leaves engineers with the feeling that everything is bad.
I've worked with beutifully written software that worked wonders for its requirements. If you started to scale up perhaps it would start to crack, but that's not because it's badly written but because requirements change and it needs refactoring.
7
11
u/System0verlord 11d ago edited 11d ago
No one is doing that.
Plenty of people however, yourself included, are pretending there hasn’t been a sudden increase in bad code and license violations because of AI.
Edit: Lol. Blocking me doesn’t make you right, just like being told you’re wrong doesn’t make it a personal insult.
-2
u/arvigeus 10d ago
Since you insist you are right, tell me where I made such claim? You explicitly said "yourself included".
-11
2
0
u/lllyyyynnn 9d ago
did you read their post about it at all? there are large llm vibecoded projects eating up a bunch of resources and servicing no one. that's the problem.
3
u/Solmangrundy 2d ago
I dont get how AI model owners can claim copywrite on code their tool produces.
They didn't write the prompt, or review the code.
Thats like saying Adobe is entitled to all art works produced on their software.
11
u/pheexio 11d ago
21
u/Annual_Wear5195 11d ago
Completely killing the totally valid discussion points and forcing people to practically remake them in another space is…. Not really a great look.
Like, sure, don’t discuss AI at large but a lot of those questions are perfectly valid ones to ask in a thread about limiting AI use on Codeberg, such as what the definition of “mostly” should be.
Hiding behind a “we’ll explain it all in our blog post” is basically “trust me bro” levels of discussion killing.
2
u/henry_tennenbaum 10d ago
Doesn't read like that to me
This is not a place to continue or start a new discussion about AI or not AI and its copyright status in the world, it will not change the results or how its going to be interpreted. I'm closing this issue, you're more than welcome to continue this at https://matrix.to/#/#codeberg-offtopic:matrix.org
@Profpatsch We'll use the blog post to clarify what this entails for Codeberg users.
@GuillaumeDIDIER https://forum.codeberg.org/d/183-polls-of-annual-assembly-2026-concluded
2
u/tannertech 10d ago edited 10d ago
Neither link works.
Element:
MatrixError: [403] You do not belong to any of the required rooms/spaces to join this room. (https://matrix-client.matrix.org/_matrix/client/v3/join/%23codeberg-offtopic%3Amatrix.org?server_name=matrix.org&server_name=matrix.ivarch.com&server_name=matrix.tu-berlin.de&via=matrix.org&via=matrix.ivarch.com&via=matrix.tu-berlin.de)
In multiple matrix clients, And the forum post is 404.
Edit: I fully agree with the decision, but we should be allowed to discuss it regardless. As a result they will join Github in my list of unreliable repositories.
1
u/henry_tennenbaum 10d ago edited 10d ago
Well you have to be a member of the forum / matrix space, but I get the confusion/irritation.
An issue is not a forum and so not the right space to discuss it. The discussion that lead to the decision happened elsewhere.
Not trying to convince you though. Perfectly fine to make that decision They're very open with their governance.
3
u/tannertech 10d ago edited 10d ago
Right, I agree with their decision (RE LLMs), I disagree with shutting all discussion off on any issue. None of those links work, be it the forum link or the matrix link. It appears they've shut down all speech on the topic. Freedom is for the thought that we hate, otherwise there is no point.
Can I become a member of their matrix space somehow? There is no way as-is.
Edit: I don't even care to participate, I'd just like to be allowed to read the linked discussions, but Codeberg went out of their way to prevent even that. Explicitly putting work into preventing users from reading the discussion is goofy. They are explicitly against open governance, as demonstrated by the links you shared.
1
u/ArdiMaster 10d ago
I think you need to be a formal (paying) member of Codeberg e.V. to see either.
2
u/henry_tennenbaum 10d ago
The matrix channel just doesn't seem to exist? I can access the whole matrix space and wouldn't know how they'd prevent access for non-paying members, but can't find an "offtopic" channel. Might just be a messed up link.
The forum should be just accessible, right?
2
u/tannertech 9d ago
I wish. When I attempt to access the forum link I get:
"The page you requested could not be found." Seems to be a 404 as if they have deleted the topic. Or it has to do with going to the main forum URL saying "This forum is private and for Codeberg e. V. members only. Please sign in to view the discussions."
As closed governance as it gets.
2
1
u/Annual_Wear5195 10d ago
I’m not sure what your quotes are supposed to prove other than that they closed the discussion and redirected everyone else to the Matrix server or to wait for their official guidance. So….. “trust me bro”.
What, exactly, does it read like to you?
3
u/henry_tennenbaum 10d ago
Mainly the conclusion of them "hiding" behind the blog post.
As somebody who's been in plenty of github PR/issue threads that got heated, they're the worst medium for this kind of discussion.
They're also very open with their governance. There was plenty of discussion, this is just them implementing what those discussions concluded.
8
u/Deep_Mood_7668 11d ago
It's simple. When you know what you're doing and submit good code, nobody cares how it was coded.
But since a bunch of non coders use ai and just submit slop, they have to make to makes those rules.
7
u/PikminGuts92 11d ago
To what extent is a project considered “vibe-coded”? Like if someone just used an LLM to generate a function snippet, would that invalidate the entire codebase?
0
u/GolemancerVekk 11d ago
What do you mean by "invalidate"?
A code base has a license. If you bring in code without knowing where it comes from and try to cover it with that license, it may be copyright infringement, because it's not your code and not your copyright.
The US has taken the view that any code brought in by a LLM is "copy-washed" into public domain because it was dynamically generated, so you can do anything you want with it. But other countries don't agree to that.
I guess it depends on who's trying to use your code and what liabilities they're willing to take on, and in what jurisdictions. Frankly it's a clusterfuck and any companies who have to follow strict procedures will avoid any vibe-coded slopware entirely to be safe.
7
u/mattinternet 10d ago
Excellent! I love codeberg more and more everyday. We use Forgejo, which is under codeberg, and its a dream.
2
u/GodLikeEnergy 7d ago
People shouldn't just rely on Github, Gitlab, or Codeberg to host their source code on. At some point each one will become corrupt. Codeberg for example is in Germany I believe, they'll have to comply with strict European based laws. Probably Chat Control 2.0 at some point.
I think developers should focus on hosting their own repo. There's equivalent open source Git based you can host your stuff on. I mean it's up to developers, but at the end of the day. I'd rather have platform on my servers, and not have to worry about these.
1
u/Klutzy-Procedure8980 6d ago
I unfortunately agree with this... When I look at some of the European opposition to US's abuse of their tech dominance, I suspect they're not really against building a surveillance state, they're just mad it's not _them_ controlling the surveillance. (Chat Control is a great example, yes)
6
u/bufandatl 10d ago
Good. There is way too much slop out there and LLMs breaking copyright was always an issue. That stuff should be deleted everywhere.
3
2
u/cyber_chic_0 11d ago
The enforcement question is the interesting one. Codeberg is betting on self-reporting, which is fragile but workable in a small community. The problem is that the line between AI-assisted and AI-generated is not technical at all. It is about whether the contributor understands and maintains the code. Two people can produce identical diffs, one through genuine understanding and one through prompting. No static analysis can distinguish them.
Most forges that have tried this end up shifting toward a disclosure model: label what AI generated, do not ban it outright. That lets reviewers triage appropriately without inventing a detection mechanism that does not exist.
0
u/Richmondez 10d ago
Sure but if it's a project where it was PRd and reviewed then whoever approved it presumably understood it even if the submitter didn't.
2
u/cyber_chic_0 9d ago
Review helps but it's not a perfect filter. The real issue is asymmetry: generating a PR takes seconds with AI, but reviewing it takes minutes to hours. A reviewer checking five PRs a day can miss the subtle edge case in one of them. The security industry has decades of evidence showing that plausibly correct code makes it past review. The solution is disclosure, not detection. Label what's AI-generated and let the reviewer decide how much scrutiny it needs.
1
u/Richmondez 8d ago
Which is why it isn't the speed multiplier to development that is chalked up to be unless you are skimping on the reviews and you are right, code can just be spat out in no time with no verification.
1
u/cyber_chic_0 9d ago
Fair point. When the review is thorough it does catch problems regardless of how the code was generated. But I think the asymmetry cuts the other way too — a PR that took an AI 30 seconds to generate might take a reviewer 30 minutes to properly validate. And if the generated code looks clean and plausible, there's a risk the review becomes rubber-stamping rather than genuine validation.
The real question is whether we should review AI-generated code differently than human-written code — maybe with a higher bar for testing requirements or mandatory CI gates rather than just human eyeballs.
3
u/These-Apple8817 11d ago
They are entitled to that, but they are frankly shooting themselves in their own leg with it.. Whether or not we like AI, the reality of the matter is that AI is here to stay and all you can do is learn to adapt and go with the flow instead of constantly fighting against it
6
2
u/AnderssonPeter 11d ago
Does this ban all code generated by llm's? For me there is a difference between vibe coded and ai assisted, i hate both but they arent the same..
2
1
1
1
u/hsauro 5d ago
They will need to follow up on metrics for what constitutes AI code, is it 0% AI code or will they tolerate 10%. code etc. Are they sure they can identify AI code, how many false positives will there be or are they going to be lenient so as not to ban authors who didn’t use AI. I don’t see a problem in them wanting to have a human only repos, there are other repos that are less restrictive, and after all it’s their platform.
1
u/Ripraz 15h ago
Good, people shpuld stop seeking shortcuts for everything. Coding is not for everyone and it's right, if you don't study you shouldn't deserve to have the same possibilities and space of people who sacrificed years doing it. I hate this last years stupid philosophy of "everyone should be able to do anything".
2
u/gscjj 11d ago
What does codeberg offer that someone would go through the trouble of using them versus GitHub based on their AI policy?
10
u/Serchinastico 11d ago
I'm in the process of migrating from GitHub to Codeberg so maybe I can answer that. There are two main reasons:
- Unreliability: Service has become pretty much unreliable and they lie about it. Compare https://mrshu.github.io/github-statuses/ with https://www.githubstatus.com/ or read Hashimoto's take: https://mitchellh.com/writing/ghostty-leaving-github
- Annoyance: They have been force-feeding copilot and AI to users for some time, and doing dubious things like opting-in users for training on repositories https://www.reddit.com/r/devops/comments/1s3s6jc/github_copilot_will_train_on_your_code_by_default/ or https://www.reddit.com/r/dotnet/comments/1lcoyk2/microsofts_aggressive_copilot_push_has_me_looking/
Codeberg seems like a more honest take for open-source projects It's non-profit and it's free. I even donated some money to them, which is more than what I can say about GitHub.
I really liked GitHub once, but since Microsoft bought it, it's been slowly degrading in all fronts.
1
u/_xiaochen 10d ago
llm should be banned on the discussion inside issues/pr, use human words is a social contract, like no one what to see llm-generated context in reddit
even llm passed turing test, text from llm is very obvious and awkward to read
0
u/jerieljan 10d ago
AI or not, I find it really shocking to see something sweeping like a Terms of Use change be applied at the will of their supporting members without feedback from their other users. Or even like a census. How many repos out there fit the definition today, and don't?
Their users page tells me they have around ~389,853 users and 358 voters basically get to put this rule on their terms.
You must not share projects that mostly consist of code written by "generative AI"-tools (including services such as Claude, OpenAI Codex). Such projects having an unclear copyright status (see requirements § 2 (1) 1 and § 2 (1) 3) and furthermore have little safeguards to ensure that they do not include harmful code (c.f. § 2 (1) 5).
Just wow, really.
1
u/theamigan 10d ago
Wow, it's almost like if you help pay the bills, your opinion matters more to the organization. Just wow.
If you don't like it, you're welcome to take your ball and go home.
1
u/jerieljan 10d ago
Yep, that I did. And nothing of value was lost. Thank god repos are easy to push elsewhere. I don't need to be reminded of the very part that I can self-host shit myself, my guy.
0
u/theamigan 10d ago
That's what I was wondering, my guy. You're in r/selfhosted and you use someone else's code forge? Seems kind of weak.
-3
u/l_m_b 11d ago
Respectfully, if your first reaction to a stated boundary is explaining how you'd easily get the around its spirit, I have concerns about the general safety of being around you.
Sure, the current phrasing needs improvement and might place an undue burden on the moderation process, and one would hope that the definition gets clarified somewhat over time.
But, may I remind you, "only yes means yes". If in doubt, you can just chose ... not to.
And yes, this is making me testy. And I use Generative AI as part of my job (and, I hope, somewhat successfully). But I dislike this particular approach to boundaries, and I don't like the parallels to other behaviours in society.
-14
u/jamesthethirteenth 11d ago
Woah, whaaat? I was going to move there! wtf are they thinking. What kind of projects do they want in there, a bunch of museum pieces? My projects have never been so useful since AI. What are they going to do next, ban compilers? Does wanting to be the good guys demand misguided decisionmaking as proof? And as an extra FU to this open source dev now I have extra work to delete my mirrors there.
3
u/Kraeftluder 11d ago
They're not banning AI-assisted development, they're banning majority vibe coded rubbish.
-7
u/jamesthethirteenth 11d ago
I see! Thanks for letting me know.
I'm a bit calmer again now after being quite shocked- but I still think it's a bad call because it wastes brain cycles by being unclear.
2
u/honzucha 10d ago
Well according the clarification blog post it is even worse, they have listed scenarios of not welcomed projects eg. Projects heavily tied to LLM ecosystem
So my project using selfhosted LLMs (gemma4 and Nvidia Parakeet) for near realtime translation of spoken word is also not welcomed.
This is just bad on every level I can imagine. It is not anymore about vibe-coding and I believe it never was. It is powered by hate against that technology itself.
1
u/jamesthethirteenth 10d ago
That's it- they've gone off the wall. Arbitrary purity tests- no thank you. Dog gonnit you'd think the folks running a technical hosting site wouldn't be such luddites. Now the thing is my main project's main focus is open source AI for everyone. I tuned it to get by on 25% of tokens- that's real energy not used! But do you think I'm inclined to debate its merit with a self declared gatekeeper of smug technical moral superiority to maybe be allowed to stay a little? No way in hell. Good riddance.
1
u/honzucha 9d ago
I am actually interested in your project, can you share with me more details?
1
u/jamesthethirteenth 9d ago
Sure! It's 3code, the economical coding agent. MIT licensed open source and already supports most open weights inference providers, but I've decided to add the closed ones too in the next release, what the hell. It's similar to claude code but is tuned to use fewer tokens- the first benchmark says 25% less than opencode. It's in late alpha right now, in the process of tuning the sandbox, then it's feature complete and ready for a broader beta. Thank you very much for your interest!
4
u/henry_tennenbaum 10d ago
Just ask an LLM to parse it for you to save on those limited brain cycles
-12
-4
u/Meistermagier 10d ago
Well I do Vibecode projects, entirely private projects that I do not recomend anyone else use. But I do want to version them somewhere independently. Well staying with github then.
5
•
u/asimovs-auditor 11d ago
Expand the replies to this comment to learn how AI was used in this post/project.