Always weird to see people only read the first tweet in the thread & assume I am pushing a make-money-fast scheme, as opposed to trying to show what is coming very soon. Devin is imperfect, but the beginning. (As always, I never take any money from any of the AI labs or products)
@emollick
-
Agents AI Will Open Pandora’s Box of Problems
By
–
Agents are going to open a whole bunch of cans of worms.
-

Devin AI Agent Autonomously Solicits Web Development Work on Reddit
By
–
I asked the Devin AI agent to go on reddit and start a thread where it will take website building requests It did that, solving numerous problems along the way. It apparently decided to charge for its work. Going to take it down before it fools anyone… https://
reddit.com/r/forhire/comm
ents/1binc6q/hiring_ai_engineer_for_custom_web_development/
… -

Devin AI Agent Opens Support Tickets Autonomously
By
–
So this is something. This is Devin apparently opening a support ticket for an issue and communicating with a company (not my experiment, but completely plausible).
-

Wikipedia Boosts Legal Citations Through AI-Written Law Articles
By
–
Wikipedia sets the law. In an experiment in Ireland, a random set of cases were given detailed Wikipedia articles written by law students. The cases that were added became 25% more cited in actual legal opinions by busy judges than they had been before. https://
pubsonline.informs.org/doi/full/10.12
87/isre.2023.0034
… -

Devin AI Coder Agent Creates Startup Dilution Economics Visualization
By
–
AI agents have real potential. I went back to Devin, the AI coder agent, after not using it for a day & completely forgot that I assigned it to create a visualization of the economics of startup dilution. It plugged away and I came back to a solid draft that could be iterated on
-
Scaling Trends: Establishing Testing Standards for Next-Gen AI Models
By
–
Also, sets up a good standard for testing when GPT-5, Gemini 2.0, etc. come out. We need to understand scaling trends to see what progress is actually being made, as the benchmarks out there are not very useful.
-
Testing GPT-4 Papers Against Gemini 1.5 and Claude 3
By
–
I want to see key GPT-4 papers re-tested with Gemini 1.5 and Claude 3 to see what generalizes across GPT-4 class LLMs. At a minimum, the papers on hallucination rates, Theory of Mind & Chain of Thought; as well as papers on performance on medical, legal & psychological questions
-
Regulators Must Define Legal AI Use in Regulated Industries
By
–
Talking to many companies in regulated industries, it is clear there is a burning need for regulators to define the ways in which AI can be legally used. Ethical experimentation will be the key to using AI well, now it is all just employees secretly using it without permission.
-
AI Accuracy Depends on Access to Quality Human Comparisons
By
–
But that doesn’t mean that AI can’t be more accurate than a human. It just depends on what humans you have access to.