Let’s imagine that GPT-5 is super-intelligent, but works only via chat interface, its speed is within human abilities to observe in realtime, and OpenAI doesn’t release it to public, but undergoes rigorous testing. Wouldn’t they learn valuable lessons while staying safe?
@marek_rosa
-
Training Models: Foundation for Effective Testing Processes
By
–
Because if you don’t train it, then you have nothing to test?
-
Training vs Testing Strategy for AI Models
By
–
Isn’t it better to continue training but double down on testing?
-
User Reviews Vicuna as Surprisingly Capable Language Model
By
–
I just played with Vicuna and it's a surprisingly good LLM! https://t.co/7VQpQKKTl6
— Marek Rosa | European🇪🇺 | South African🇿🇦 (@marek_rosa) 2 avril 2023I just played with Vicuna and it's a surprisingly good LLM!
-
Generated Code Replacement and Program Decomposition Strategy
By
–
Aha ok. As far as I know we are replacing the whole generated code, not parts of it. Decomposition of the program is one of our next steps, but I also think it’s much harder.
-
Self-modifying AI agent code recursive implications
By
–
No, I mean if the code of this agent/script is being recursively modified by itself?
-
Is AI Really Recursively Self-Improving Its Own Code?
By
–
Is it really recursively self-improving? So not just generating code, but also modifying its own code?
-
Agent AI Corrects Generated Programs Using Test Feedback
By
–
By patching you mean how the agent corrects the generated program? We get the feedback from the tests and feed it back to LLM, to make an intelligent decision on what needs to get fixed.
-

GoodAI Autonomous Agents Design Test Programs Autonomously
By
–
In GoodAI, we are working on autonomous agents powered by LLMs. This is one of those projects where the agent gets an initial goal (to build a program) and then autonomously designs tests, runs them, and proposes corrections until the program works. Another use case for
-
Current AI Doom Drama: Key Thoughts and Perspectives
By
–
This is what I think (for those who follow the current AI doom drama)