FINERACT-2800: add ai policy - #6366
Conversation
2f742aa to
5dc1153
Compare
- link to ASF guidance - include my language on this topic from https://lists.apache.org/thread/blqqopc0t0m1yc3fq367zhtrq3rkj4w2
- move RAT instructions up, into "Developer How To's" - move building documentation instructions up, into "Developer How To's" - move AI policy down, into "How We Code"
I missed this one during FINERACT-2734
current source tarballs *do* include the gradle wrapper (maybe old ones did not)
5dc1153 to
d79858f
Compare
|
@meonkeys the bare minimum I think we should do is actually require the disclosure of AI tool usage. We will not do ourselves a favor if we might be later required to sort this out (if/when we receive complaints). Using AI tools to detect if a file was generated by AI tools is very unreliable and gives at best some percentage values that you still have to interpret yourself; I see currently no way to automate this. Disclosure should be anyway easy enough: we have to write these commit messages anyway and we already have all those checklists in the current PR templates, so we could extend the checklist there. Right now I'm only aware of Anthropic watermarking their generated results (no idea how it works nor how it could be used to detect generated code automatically). In the end this is then only one provider anyway, others don't do this, maybe later some standard will emerge eventually. If that happens then we could probably forget the manual disclosure and automate this process. Until then personally I'll mark every PR/commit with a note which provider/harness was used if parts of the code were created with the help of an AI tool. At the moment the only scenario that I could imagine this to happen in my workflow is if one day an AI based security scanner shows up that auto-suggests the fixes. Right now I don't see this happening (I recently tested a provider... not worth the money... at all). Most likely though I would use AI to kickstart documentation, but would still mark it even if I think that I've changed it considerably. And in general I find all of the providers in various degrees useful for discovery/research. So, if we feel that as a community not disclosing generated code puts us strategically in any better position and this conforms with he Apache license and the current responsible AI usage rules then be it. My concerns are still:
I have more, but for the sake of coming to an end here I'll stop. I have a feeling that the moment for a proper (pros and cons) discussion has already passed (see "inevitability"... what a word); not expecting all those points to be answered, but wanted to post them anyway. I really hope that these tools help everyone create better and more robust code (not necessarily more). I'll try with a more "classic" approach, but glad to change my mind. For now though my personal process won't change. I'm reviewing these tools once in a while and if I see a sensible opportunity I might change my setup (still with disclosure), but only if I feel more comfortable with the consequences (you can't control everything... I know). In all, just my 2 cents. |
based on ongoing discussion with @Aman-Mittal and @vidakovic
8a4f208 to
81da32f
Compare
|
@vidakovic thanks for adding your thoughts here! I made changes! Please see 81da32f.
You're leaning "disclosure is required" and @Aman-Mittal is leaning "optional", so I'll split the difference and make a "recommendation". Regarding "disclosure is required"... why? This can be a side-discussion or whatever, I'm just curious. From https://www.apache.org/legal/generative-tooling.html (as of today) I'm getting "it's a good idea" but not "gen AI use MUST be disclosed". I want to adopt ASF's policies/recommendation around AI use as they're doing their homework on it (and I'm not).
Sure, but--just a side note--I think the checklist is dumb. 98% of the time it is just rubber-stamped. Maybe we require it for new/drive-by contributors, but otherwise it seems like useless red tape. Nevertheless, I added a checkbox in 81da32f.
Personally I'm fine with whatever channel fits the use cases of expected response time, privacy, and verbosity, as long as whatever is discussed is logged and searchable forever and copius cross-references are provided. I question your choice to repeat most of your concerns about AI here... This PR is really just about a few salient lines in
I disagree. I do think the mailing list back and forth wasn't really helping us move forward. Maybe call a meeting instead? I'll gladly attend. For now we just need something useful in our contributor guidelines. We'll iterate and improve it. |
TianHengZhuang
left a comment
There was a problem hiding this comment.
Nice additions! A few observations:
- AI Policy checkbox: Good initiative for Apache projects post-LLM era.
- zulu21→zulu25: Ensure this is aligned with other docs and workflows that reference zulu21 — should those be updated together?
- RAT tool & doc build instructions: Useful additions for contributors.
- WIP: Will the AI Policy section in
CONTRIBUTING.mditself be added as part of this PR? Worth verifying the link anchor#ai-policyis valid before merging.
Overall a solid policy contribution. ✅
|
@TianHengZhuang wrote:
uh, was that an AI review? Please review that review for accuracy. I'm seeing several issues with it. (@vidakovic the irony here is killing me. 🫣) |
@meonkeys Haha no Meonkeys, I just care about wording. If there's something inaccurate, feel free to point it out — happy to discuss |
|
@TianHengZhuang please take the time to carefully review the commits, diff, and commit log messages. |
FINERACT-2800
dev list discussion