Your RL framework or your RL project has no chance to be a game changer if it's written by Claude or ChatGPT
I'm kinda tired of vibe coded repos that are presented here as some kind of archievement.
Do you know how your code works? Have you ever written PPO from scratch? Do you know about policy gradients?
Did you filter your data?
If not, please dont spam - except if you have questions. Then we are happy to help newcomers oder advanced learners.