I don’t really know where to put this topic, so feel free to move it somewhere more appropriate.
We’re already dealing with some users submitting “AI” generated bug reports, issues etc. I also suspect that we deal with at least partly “AI” generated content/claims in other dialogue, even when it’s not a pure copy/paste.
It might be impossible to properly tell what is what, and I expect this to be a problem that will only grow larger, but I think it’s time to consider having some stated policy on the topic.
The problem is that while it’s very convenient for the user to let an “AI” file their bug report, it’s often both verbose and full of errors and erroneous claims, mixed with genuine information and correct conclusions. Unentangling this mess is hard and frustrating work, and very demotivating for people trying to help others.
If you can’t even trust that the other part is at least trying, to the best of their ability, to provide correct information, the whole concept of “helping others” quickly become untenable, or at least that’s how I feel it.
While it wouldn’t solve the problem, I think it would be helpful if there were rules that stated to which extent users are allowed to submit “AI” generated bug reports or problem descriptions. That would at least allow us to refer to those “rules” instead of having to face it “on an individual basis”. I have already had numerous occasions where I’ve had to state that I’m not willing to engage/help because of “AI” generated input that makes me waste a lot of time chasing hallucination and trying to separate fact from fiction. It always feels like I’m “uncooperative” when I have to do this - the alternative is what I think many does - simply to not respond. But, neither is optimal.
The rules for each repo is ultimately set by the maintainers of the repo. Theoretically a policy could be raised up to the AC level but I don’t know that we are at this level yet. You might open a discussion issue on the repos where this is occuring the most (core and webui probably, maybe addons) and see if the maintainers can come up with a policy that makes sense for them.
If I were king I would set the policy as follows:
Use of AI in responses must be disclosed so the developers can assess how reasonable the information being posted is.
Use of AI without disclosure which is later discovered may result in the issue summarily being closed with the option to reopen if the original issue filer materially participates in the thread and discloses when it it them and when it is the AI talking.
Use of AI in a PR or code is OK but it must be thoroughly tested before it’s submitted as a PR. (I’m sure this could be written better and I don’t think it covers everything.)
Repeated violations of the repo policies can result in becoming blocked from opening new issues and PRs. (Assuming this isn;t already a rule.)
I really understand why Flatpak has just banned AI PRs entirely. It’s not just the poor quality of the code and the interactions, but the people who are doing this tend to be entitled and rude. I definitely do not want to see that become a problem for OH.
I tend to disagree with your #1 and #2
Our (= all maintainers’) goal should be to have no AI reports without a big button disclaimer on the front because, as @Nadahar absolutely correct has put it,
For that to happen, we need to have an OH wide policy, put up by the arch council and enforce it (eventually even sanction).
It’s up to every maintainer what (s)he makes of it (such as to create another issue to contain valid parts), but without a strictly enforced and comprehensive, congruent policy, we’ll end up discussing if to apply rules and which ones for each and every issue in question.
The only authority the AC council really has is if:
there is a desagreement among the maintainers of a given repo that they cannot resolve among themselves
there is a disagreement between mainters of different repos that they cannot resolve among themselves
We have no such controvercy here yet so the AC can’t get involved. The individual repos need to give it a try first I think.
Of course, any maintainer (I think any member of the openHAB project on github actually) can raise an issue to the AC, but based on the rules of the AC I don’t see how they have authority to impose it. From the AC’s about:
The purpose of the AC is to be an authority to appeal to when there is an impasse among the developers and maintainers of the various parts of OH or to help maintainers from different parts of OH to come to a decision on how to address a cross cutting change. It is not intended to be a place where dictates are are handed down from on high. Team Discussions: openhab/openhab-distro
The AC can help negotiate an OH wide policy but they cannot impose a policy from on high nor can they enforce it.
I’m not after a “legally binding” policy, but an “OH wide” guideline, so that everybody don’t have to reinvent the wheel, and for consistency for users as well. I’m thinking that it should apply to forum posts as well as GitHub issues, the core problem is the same - if somebody should use their time trying to figure out what’s causing the problem, they should at least know that the information they are provided is genuine.
So, if you want to do it the “judicial way”, maybe what I’m asking for is a guideline that various repos and the forum can opt to follow. Those of us that don’t enjoying wasting time chasing down hallucinations could then opt not to participate in those areas that choose not to follow the guideline.
Regarding your points, I agree with #1. I don’t think it should be outright banned, but it must be clearly marked so that there’s no doubt what is genuine and what is “AI” generated content. I’d say that #2 might need a slightly stronger tone, but I agree that it shouldn’t automatically mean that it’s closed - but that might be a necessary measure if people aren’t willing to comply disclosing.
Regarding #3, I think I pretty much disagree, the exception is very simple PRs that update dependencies etc., but those should be initiated by maintainers, not by users/contributors. This is outside the scope of what I meant to deal with here though, this was only meant to be about the communication itself.
I fear that #4 is necessary, but shouldn’t be handed out eagerly. People must have the chance to understand what they do wrong, and still keep doing it before I think a ban is appropriate.
I absolutely understand what Flatpak has done, and I think we’ve only seen the start. The problem is that it can be hard to tell, so I’m fearing that this can lead to overly harsh “punishment” to try to compensate - as humans usually do when they have a problem that they can’t really get a grip on.
But, at least having a guideline on the topic is necessary, because otherwise you can’t really blame people for doing it. You must try to explain it to each individual and appeal to their understanding of what they expose others too - a very hard thing to achieve with some people (the entitled/rude group are pretty much immune to reason and considering the situation for others).
I think just ignoring the issue guarantees that it will just be a growing problem.
Back when it was established @kai was very concerned and very deliberately wanted the AC to be something like a “supreme court” rather than a dictator. And that’s mostly how it’s operated since then.
It could be more proactive but that would go against it’s original intent and I personally, speaking as a member of the AC, would feel uncomfortable if it were to do so unless there were unanimous consent among all the maintainers (or at least a super-majority) of all the OH repos to give it that authority. But then be careful what you ask for because this won’t be the only issue the AC will have something to weigh in on.
I’m ambivalent on this one. For one, I don’t think it’s a problem yet on the forum. I can definitely see it become a problem but so far it’s pretty obvious when AI is involved and it hasn’t been too much of an issue yet (or if it is I’ve missed it). I defintely can get behind a “use of AI must be disclosed” policy but it will be hard to police. Even I don’t read every post made to the forum and there are a a few users I’ve added to my ignore list (mainly bcause of past behavior).
I also don’t want the forum to become too quick to accuse users of posting AI slop. Assuming users are telling the truth (I’ve no reason to believe otherwise) I’ve mistakenly accused users of posting AI generated weird code when it was in fact their own. If we start pointing our fingers at everyone I fear it will make the forum less welcoming and friendly, and those are one of OH’s greatest strengths I think.
Sure. Any member of the openhab project on openHAB can start a new thread for the AC. Anyone can open a discussion thread on any openHAB repo. We don’t need consensus here do get something moving. I think starting with the repos is a better approach but I can’t stop anyone from moving forward and going straight to the AC (and I wouldn’t do so if I could).
My intent with 3 was that any AI submitted PR (I agree excluding dependency upgrades) should be expected to have been tested by the submitter at least as much as we expect a manually created PR to have been. I’ve seen on other repos AI submitted PRs that won’t even compile and certainly don’t pass the unit and integration tests.
That’s why I said “repeated”. A pattern of behavior must be established and the user had plenty of communication informing them of where they are going wrong and how to correct the problem. And I would not make the bans permament either.
Again, enforcement isn’t my primary concern, I think that stating that this “isn’t OK” is worth something in itself. I’ve seen it on the forum. It’s not that the users start with AI content, but rather that when you ask them questions they might not know how to answer, some tend to just feed it to some “AI” instead of asking for clarification about what they don’t understand or know how to do. And, then they either paste the “AI” answer, or refer to parts of it like it’s some kind of fact, and we can go many rounds before it eventually comes out that this was never a fact, it was an “AI” claim.
To me, the most important aspect of this is to make it clear that dumping “AI” garbage whenever you don’t understand something, isn’t “good etiquette”, it isn’t respectful, and those that you communicate with aren’t likely to appreciate it when/if they figure out that this is what it is. When nothing is stated on the topic, we rely on each individual to realize this themselves, and I think that the chance that those that never try to solve problems themselves realize this on their own, is slim.
I agree with this concern, I have a “writing style” that some think looks like “AI” content, so I’ve already experienced being rejected because of this. But, the solution can’t be to take no stand, because individuals will be put of by it regardless. Saying that “if you post “AI” content, make sure it’s clearly marked as such”, doesn’t really impact the threshold for “accusing” people of doing it as far as I can tell.
I just raised this is a suggestion, something to think about. I’ll not proceed with any “formal requests”, that’s not really my point. I just want people to think about it, to figure out how it should be handled, instead of just letting it grow to a very tense issue with lots of strong emotions. I already “declare” that I’m not willing to spend time on such content, I can just keep doing this for my own sake, I just fear that this will become something that grows increasingly tense and sour.
It’s worse than that in my opinion. It’s relatively easy for an “AI” to make sure that it passes tests, it can just tweak the code until it does. That doesn’t make the code correct, or even sane. Tests are nowhere near being something that we can rely on to assure correct behavior, I don’t know if that’s at all achievable, but it certainly isn’t reality. If passing the tests is your goal, it’s usually pretty easy to tweak the code to satisfy the tests (like VW did with their diesel emissions).
I’m not sure that I’m interested in “AI” generated code at all, in most circumstances. It’s just not worth the hassle to try to discover and correct all the “logical flaws” it’s done along the way. If you don’t know how to solve something, have it generate suggestions and take inspiration from some of this if useful, but don’t actually use that code, is my take.
I overall agree with the proposed guidelines above and also think we need a policy (coming from the GitHub issue that @Nadahar posted above).
Wrt who is the “authority” to put a guideline in place, I don’t think we need the AC. IMO it should be enough to start a openHAB Maintainers Discussion, propose a guideline and ask for feedback and votes there.
Assuming agreement on open-source application complexity, AI might be what finally draws a potential person with Home Automation interest to try OH. I’d go lightly on the user side. Maybe some verbiage in the pinned Ask a good question post. An AI assisted post might be all a new user can do. It could just be off because they didn’t describe the right issue or the right issue correctly to the AI.
For maintainers (assumed different level of skill). I’m okay with whatever but likewise wouldn’t want to discourage new contributors. There are a lot of rules already. Obviously (to me anyway), any code needs to be tested, compile and pass any tests.
Github issues are in between. From my own perspective it was a long time (after using the forum exclusively), so more advanced user IMO. There are already templates, but maybe some verbiage there too.
I must say that I don’t quite understand what you’re saying. Are you saying that you think that “AI” can somehow help recruit new users? If an “AI” assisted post is all a user can do, what would he/she have done a few years ago?
This is how I see it: The user has a problem/challenge or question. The user must manage to express that challenge or question in some way, or the “AI” can’t “do” anything either. So, it’s between 1) The user expressing this directly to us 2) The user expressing it to an “AI” that will “fill in” a lot of details that might be spot on, might be slightly off, or might be completely off. By presenting the “AI” generated content to us, that originally expressed to us is mixed with all kind of speculation and hallucination. With 100% confidence and conviction as always. We have no way to know what parts are what, what is actual information coming from the user, and what is… fiction. It’s simply much better for everybody if we don’t have the “AI” obfuscate the real information with irrelevant or wrong information. If we want to use an “AI” to try to analyze a situation to see if it comes up with something worth examining, we can do that after receiving the issue/question from the user. That way, we know what is what, we can differentiate between “actual information” and information which might or might not be true. It’s important to know what is what when you’re trying to solve a problem.
Are you saying that by not allowing “AI” PRs, we would discourage new contributors? To me, there’s an ocean between a (code) contributor and a “vibe coder”. I personally think that “discouraging vibe coders” is a good thing, the effort you have to put in to make something worthwhile out of such a starting point just isn’t worth it.
You can’t put too much trust in tests. I know that there are people that swear to tests with a religious conviction, but my claim is that to actually write tests that cover all the bases is a mammoth task, and in some cases simply isn’t possible. It depends heavily on the code, some types of code can be covered pretty effectively with tests, others I feel that the tests are of marginal value. There are so many things that can go wrong that you just can’t capture, at least not without simulating a huge system.
Let’s not forget that tests are usually written by the author of the code. It’s not like somebody writes the tests first, and then somebody else makes the code to fulfill the tests. They are, in all normal circumstances, written after the code. The tests exist mostly to prevent later modifications of the code from breaking what’s there, and they often don’t do a very good job at that either.
I often find that writing reasonably proper tests for a piece of code can take 3-4 as long as writing the code itself. And then you still have a lot of “loopholes”. If you want to make the tests really foolproof, I think the time required to write the quickly approaches infinity.
“AI” are also in a particularly good spot when it comes to “fooling” tests. They just run them, tweak and reiterate until the tests don’t fail. That’s not a good way to identify structural problems with the code or logic, that’s just a way to tweak behavior until it satisfies the tests, which might mean that the code is much worse at performing the task it’s meant to perform than before it was “tweaked” to fit the tests.
It’s similar with reviewing code. You can review code to see if what happens looks reasonable or some obvious mistakes take place. But, if you really want to follow every nook and cranny, every piece of logic, then you must reconstruct the whole thing in your head. That is significantly more work than it would have been to just write the code yourself in the first place, because not only must you figure out how the problem should be solved, you must also “reverse engineer” that thought process of the author and compare that to your own idea of how it should be solved.
It’s simply just a limit to how much you can “expect” from a review. There has to be some level of trust, some assumption of intelligence and good intent, or the task becomes too big. Reviewing is already the biggest bottleneck in OH, it won’t help the situation at all if you allow flooding the “review queue” with “AI slob” that might not make any sense at all in the end. It might be structurally or logically invalid, but it still (often) requires a significant effort just to analyze it enough to draw that conclusion.
Nobody is saying that they can’t use “AI” themselves, or even that they can’t post it - but it must be made very clear what part of the text is “AI” and what are their own words, so we don’t have to play the guessing game about what is what.
The issues created so far are strictly related to GitHub. The forum is always going to be less strict than GitHub when it comes to these things. But the biggest thing is disclosure. If AI is in use, tell us. That’s the biggest thing. If we know then we can address on our own whether it’s worth our time to engage or not. And we can avoid a situation like occurred on GitHub where there was 6-12 posts before it was revealed that. an AI agent was making all the replies independently and the person was just monitoring.
That’s the sort of situation we want to avoid. It’s not the use of AI that’s the problem per se, but that the AI was making assertions about things and it was being wrong in a way that caused a lot of extra work for the maintainers because a human user wouldn’t have made such incorrect assertions. If the maintainers knew it was AI, they would have been more skeptical of the assertions made, or they could have decided it was a waist of time to engage at all.
Far too much time is waisted chasing hallucinations from AI. The absolute least thing that can be done is to disclose when AI is involved so they at least know such hallucinations are possible.
But I would vote that it has to be disclosed in the forum too if the content of a post or parts of it are AI-generated. Maybe with a standardized marking e.g. [AI generated], [AI structured]… before the corresponding part.
If I’m about to answer a question in the forum I definitely want to know if I’m corresponding with an AI.
I think that’s somewhat of an oversimplification. First, it’s my impression that many people don’t read anything that’s more a line or two, regardless of what the source is. Second, it’s not necessarily that you don’t want to read it, but if you read it, you read it knowing that the things stated in the text may not be factual at all.
You could argue that people can’t be trusted and might post unfactual information too, but I think there’s a “huge leap” between the amount of “untruths” posted by humans and that of “AI”.