Can AI build an entire software project from scratch, not just fix one bug? Researchers at Meta FAIR, Stanford, and Harvard introduce ProgramBench. This benchmark tests if language-model agents can take a program’s documentation and build a full codebase that behaves
GENERATIVE AI
-

Using LLMs for Legacy Mainframe Program Upgrades
By
–
Mainframe Pros on Legacy Program Upgrade Using LLM! #BigData #Analytics #DataScience #AI #MachineLearning #NLProc #LLM #IoT #IIoT #PyTorch #Python #RStats #TensorFlow #Java #JavaScript #ReactJS #GoLang #CloudComputing #Serverless #DataScientist #Linux #Programming #Coding
-
AI agent one-shots video with multi-tool workflow
By
–
just tried this out and it one-shotted* this video: "before the agent does anything"
— Yohei (@yoheinakajima) 13 mai 2026
*i generated the narrative using chatgpt and used that as a prompt. featuring: @e2b @runanywhereai @composio @mem0ai @firecrawl @browser_use @agentmail @covenantlabsai
some thoughts:
– i… https://t.co/X3RgWs9MOk pic.twitter.com/xjmXZF11HPjust tried this out and it one-shotted* this video: "before the agent does anything" *i generated the narrative using chatgpt and used that as a prompt. featuring: @e2b @runanywhereai @composio @mem0ai @firecrawl @browser_use @agentmail @covenantlabsai some thoughts:
– i -
OpenAI GPT-5.6 close, question about Anthropic Sonnet 4.8
By
–
Seriously, OpenAI is on a run. GPT-5.6 very close. Where is even sonnet 4.8 @AnthropicAI ?
-

CausalCine: Real-Time Autoregressive Generation for Multi-Shot Video Narratives
By
–
CausalCine Real-Time Autoregressive Generation for Multi-Shot Narratives
-
Stop treating AI prompts as magic spells
By
–
Stop turning prompting into magic spells (and yes, this includes random slash commands with obscure outcomes). Let this one area of working with AI not be weird. Just ask for stuff, in well-specified formats, like a manager, not a sorcerer with a bunch of incantations.
-
Development Environments Enabling End-to-End AI Agent Task Execution
By
–
Customers like Decagon, Amplitude, BILT, and Snyk use development environments to let their agents handle tasks end-to-end. Learn more:
-

Anthropic Increases Weekly Limits for Claude Code
By
–


Anthropic raise weekly limits on Claude Code by 50% until July 13! Sounds like Colossus 1 came into play!
-

Agentic Discovery for Test-Time Scaling of LLMs
By
–
LLMs Improving LLMs: Agentic Discovery for Test-Time Scaling Zheng et al.: https://
arxiv.org/abs/2605.08083 #ArtificialIntelligence #DeepLearning #AIAgents -
OpenAI Begins Internal Testing of GPT-5.6 Model
By
–
https://t.co/lxcgRqeVas pic.twitter.com/MLeGYsUxST
— Bojan Tunguz (@tunguz) 13 mai 2026SCOOP: The development cycle for GPT-5.6 is now in full swing at OpenAI. The first checkpoints of the model began testing internally over the last few days, with a release likely next month. x.com/synthwavedd/st…
