thanks! would be cool for example for openclaude traces to be uploading by default to HF in private and then people can make them public if they want to!
CODE
-

Meta Harnesses: Automated Framework Optimization for AI Tasks
By
–
Meta Harnesses is Autoresearch on steroids. Something I've been exploring recently is to get long running agents to hill climb on a verifiable task to continuously improve without my intervention. Karpathy's Autoresearch did this pretty well on specific tasks, but this weekend I tried Meta Harnesses which moves one level of abstraction up. What does Meta Harness do? Autoresearch can be used in harness like Claude Code / Codex to generate experiments to try, evaluate results, and continue looping. Meta Harness generates a harness itself that optimizes on a task or a set of task. Here, we define a harness as "a single-file Python program that modifies task-specific prompting, retrieval, memory, and orchestration logic". The idea is that LLMs are very powerful today, but to harness [pun intended] their power, you need to give it the right prompts and context. Meta Harnesses automates coming up with the right prompts and the right way to retrieve context to solve a problem. Where did this idea come from? This is from a paper from Stanford and the author of DSPy written last week. The paper shows fantastic performance on 3 tasks: text classification, math reasoning (IMO level problems) and coding (Terminal Bench 2.0), far outperforming traditional harnesses. The discovered harnesses are interesting: math for example, splits up the logic into different categories (Combinatorics, Geometry, Number Theory, Algebra) and prompts and looks at the context differently. The coding harness, amongst other things, pre-processes the tools available in the environment to save exploratory turns. When should you use and not use it? Meta Harnesses seem pretty useful for tackling a specific but wide set of problems where the result is verifiable. In contrast, when I tried it on a specific task like Chess, it arbitrarily divides the problem into separate tasks – opening, mid game, end game, and creates different approaches for each. This "works" but isn't really clean because we believe there should be one approach that does all three. It does far better on things like examinations (JEE, Gaokao) where it splits problems into categories and tackles each category with different strategies. This paper covers a pretty light version of what a harness means. In the future, we can split up tasks into harnesses that have access to specific kinds of data, specific toolchains and various models to get even better results. Overall, pretty cool applied AI approach to hillclimb a verifiable task in a specific domain with variety within the problem space.
→ View original post on X — @askalphaxiv, 2026-04-06 16:22 UTC
-
Anthropic Claude CLI Access Restrictions and Restoration Efforts
By
–
That was the plan, but then I discovered that Anthropic blocks us even when using the claude cli. Which now, apparently was a mistake? So, restoring support for that currently. It’s hard to know when the only comm is through Boris on X basically, no official statement.
-

Architect: Rapid Transformation of Use Cases into Intelligent Systems
By
–
Architect facilitates the transformation from a use case prompt to an intelligent system with a working UI in minutes. If you are looking to automate your own client-related tasks, you can explore the platform here: architect.new/ [Translated from EN to English]
→ View original post on X — @kimmonismus, 2026-04-06 16:04 UTC
-
Agentic Layer: Logic, Workflows, and Intake Data Qualification
By
–
5/ Next was the Agentic Layer. This is where the actual logic lives. Architect configured the agentic workflows to handle the thinking parts: qualifying the intake data and triggering the welcome communications based on the project type.
→ View original post on X — @kimmonismus, 2026-04-06 16:04 UTC
-

Architect generates consultant dashboard plan and wireframe
By
–
4/ The process started in Plan Mode. Architect interpreted my requirements to generate a structured plan and a wireframe for the consultant dashboard. It mapped out exactly how the data flows from the initial intake form to the tracking records.
→ View original post on X — @kimmonismus, 2026-04-06 16:04 UTC
-
One Shot Mode in Architect builds intelligent systems with UI
By
–
3/ I used the One Shot Mode in Architect. I described the specific use case in natural language and the system builder handled the rest. It did not just give me a simple prototype. It built an intelligent system with a working UI in minutes!
→ View original post on X — @kimmonismus, 2026-04-06 16:04 UTC
-
Solo consultant builds AI-powered client intake and project tracking system
By
–
1/ Engineered a client intake and project tracking system utilizing the new Architect.
— Chubby♨️ (@kimmonismus) 6 avril 2026
This is the outcome! For a solo consultant, managing the gap between a "yes" and the first kickoff meeting is usually a manual mess. I wanted to see if I could build a professional solution… pic.twitter.com/Lox5tBTQPb1/ Engineered a client intake and project tracking system utilizing the new Architect. This is the outcome! For a solo consultant, managing the gap between a "yes" and the first kickoff meeting is usually a manual mess. I wanted to see if I could build a professional solution from a single prompt. Check out the UI I built here (one shotted):
→ View original post on X — @kimmonismus, 2026-04-06 16:04 UTC
-

Agent-Powered Neuroscience Research Corpus with Open-Source LLM
By
–
Implemented @karpathy 's obsidian idea and am now letting my agent build out a corpus around my neuroscience research. Obviously still needs vetting for informational accuracy, but certainly interesting to see the web grow! All using huggingface.co/DJLougen/Harm… with @NousResearch Hermes, ~35 t/s on my 3090.
→ View original post on X — @scobleizer, 2026-04-06 15:38 UTC