I think it’s still uncertain if the current imitation learning approach will really scale to very general physical intelligence though right? But there are other approaches that might be even more promising
RESEARCH
-

OpenEnv: Evaluating Tool-Using Agents in Real-World Environments
By
–
OpenEnv in Practice: Evaluating Tool-Using Agents in Real-World Environments https://
buff.ly/viA9cQv
#AI #MachineLearning #DeepLearning #LLMs #DataScience -
Open-sourcing AI testing dataset with full transparency
By
–
We have published an extensive technical report on our methodology and we will be open-sourcing the full human testing dataset. We have always be maximally transparent about our process and our reasoning.
-
AI Systems Fall Short of Human Job Performance Standards
By
–
Virtually every human job on earth has a higher bar. These are not very high expectations for AI systems that claim to be able to do everything humans can.
-
Defining ASI: Super Intelligence Beyond Human Performance
By
–
"2+ people can do it out of an unfiltered pool of 10 people that might well be a below-average sample" is not the sign of a insurmountable challenge. It's not certainly where I would set the bar for "super intelligence". ASI is when AI is better than *every single human* — for
-
ARC-AGI Benchmark Standards and Human Performance Expectations
By
–
This is a very low bar, objectively. The claim is obviously not that 100% of humans could solve 100% of the games — that would be silly, and it wouldn't be true either of ARC-AGI1 or 2, nor of any AI benchmark that has ever been used in the field. Not even MNIST can be 100%
-
ARC-AGI-3 Environments Meet Human Feasibility Standards
By
–
To be clear, all ARC-AGI-3 environments are feasible by humans with no prior ARC-AGI-3-specific training. Our bar for feasibility is the following… Each environment was seen by 10 human testers. If 2 testers could independently clear it (successfully solving *all* levels in
-
MIRI Proposes US-China AI Superintelligence Development Halt Agreement
By
–
Speaking on behalf of MIRI TGT (not necessarily MIRI overall) We share many of the same concerns, which is why we structured our model agreement (below) the way we did. It invites broad participation, but also features mechanisms to address states which insist on operating outside of the agreement, while prioritizing the national security requirements of the US and China. So to address question 5 upfront, “should [this] be a global agreement?”: Yes! We think the US and China would be a sufficient seed to get broad participation via their network of allies, superpower status, and AI dominance. Now going point by point: “1. Assuming we achieve the desired policy goal through a bilateral US/China agreement, what would be the specific metric or objective we would say needs to be satisfied in advance? Who decides whether we have satisfied them? What if one party believes we have satisfied them but the other does not?” There are two interpretations of this question. Interpretation (1): what metric is used to determine whether the desired policy goal is being achieved? Interpretation (2): what metric is used to determine when a halt is to terminate? I’ve tried to address both below: The policy goal is to forestall the development of superintelligence long enough for other, better solutions to be realized. It is hard to say what these solutions will be in advance, as humanity is nowhere near being able to align a superintelligence. The field doesn’t have a clear path to solving that technical problem. Furthermore, solving alignment isn’t sufficient on its own, and the other thorny problems (such as concentration of power) require similar focused effort which we aren’t seeing on current timelines. The key metric we use to know if that goal is accomplished is the confidence within the leadership of the US and China that no one is advancing the frontier of AI general intelligence capabilities anywhere. This confidence is reflected by the continued willingness of these actors to participate in the agreement, and springs from a combination of restrictions/controls, transparency, verification, and intelligence gathering. It would be great if we can attain this confidence without much constraint on the beneficial uses of AI we already see today, and our agreement aims to preserve these! The agreement is not accomplishing its aims if only one of these key parties has such confidence. We have tried to accommodate the requirements we think that the USG and CCP would have, but also expect that many details would need to be ironed out through an actual negotiation and implementation effort. “2. If the goal is achieved through a bilateral US/China agreement, would we need capital controls to ensure that U.S. investors cannot fund semiconductor fabs, data centers, or AI research labs in countries other than the U.S. and China?” Yes, just like how the U.S. makes it hard for you to fund terrorists or give money to the North Korean military. “3. Would we need to revoke the passports of U.S.-based AI researchers and semiconductor engineers to prevent them leaving America to join AI-related ventures elsewhere? How else would the U.S. and China keep researchers within their borders?” There will be no shortage of technical work for talented researchers under our proposed agreement, and the best approach is for states to modify their incentives (i.e. pay them well) to act in our collective interest, in the style of efforts like the International Science and Technology Center. In 1994, the ISTC kept former Soviet nuclear researchers employed in peaceful work so that they wouldn’t sell their expertise to proliferators. We anticipate that some researchers will emigrate to non-signatories and pursue covert work, in spite of any efforts. The agreement aims to provide the US and China with sufficient confidence that these efforts will fail through a combination of compute denial, detection, and enforcement. The framing of this question seems to imply that some agreements may only aim to address AI development within the US and China, and that such development must not leave those jurisdictions. We agree that is not viable. We cover this in Article XII. “4. How should we grapple with the fact that (2) and (3) are common features of autocratic regimes? “ It doesn’t look like it takes qualitatively different “autocracy” than was required to prevent the proliferation of nuclear weapons. Limiting the development and deployment of extraordinarily dangerous technology is a feature of our American system of government which prioritizes the defense of individual life, freedom, and property. Preventing you from refining uranium in your basement and assembling a nuke in your garage is an impingement upon your freedom, but that doesn’t mean society should let you do it, and it doesn’t mean the government needs to become an autocracy to prevent it. So too with superintelligence. We charge our military and Intelligence Community with ensuring the safety and freedom of Americans against all threats. Through careful institutional design and adherence to our constitution we can avoid abuse of the power granted by our agreement. As an aside, we believe that the potential for abuse of our agreement is less than the potential for abuse of AI systems developed and employed by the government without constraint, or the potential for abuse in arrangements where the government is allowed to gatekeep access to powerful AI. Read More: An International Agreement to Prevent the Premature Creation of Artificial Superintelligence techgov.intelligence.org/res…
→ View original post on X — @esyudkowsky, 2026-03-26 23:19 UTC
-
MRM Role Critical for Gen AI Model Safety Before Launch
By
–
.
@Fidelity
's Mark Iorio proves MRM needs a role before models go live. Gen AI introduces risks traditional validation wasn't built to catch. Speed and safety aren't opposites — but you need the infrastructure for both. Learn how in our on-demand panel: https://
hubs.ly/Q048nF8q0 -

Google’s KV-Cache Optimization: TurboQuant Vector Quantization Explained
By
–
Google's new KV-cache optimization broke the DRAM stocks, but how does it work? Let's take quick a look. "TurboQuant: Online Vector Quantization with Near-optimal Distortion Rate" TurboQuant combines 2 ideas from 2 earlier lines of work: PolarQuant and Quantized