And lastly, our vision for what we believe proper test and evaluation of models really requires:
@alexandr_wang
-
New Test and Evaluation Product for Frontier Model Developers
By
–
Second, our product page for our new Test and Evaluation offering, now available to frontier model developers building on top of our work with OpenAI and other leading labs:
-

Ecosystem Approach for AI Development and Governance
By
–
Lastly, we believe this requires a whole of ecosystem approach, involving many stakeholders: – Model developers
– Enterprises and users
– Government
– Third party evaluators (like Scale) -

Hybrid Approach to AI Model Capability Evaluation Methods
By
–
For capability evaluation, our approach is hybrid again. We believe in automated methods for fast identification of performance, but most researchers do not trust fully automated evaluations for full assessments of model quality. Expert evals are a necessity for accurate evals.
-

Critical AI Risks: Bioweapons, Cyberattacks, Bias and Privacy
By
–
There are a number of very real risks: – Bioweapons
– Cyberattacks
– Bias
– Misinformation
– Privacy breaches
– Critical infrastructure This is a key problem for the future of AI -

Hybrid Approach to AI Capability Testing and Risk Evaluation
By
–
Our platform is built to both continually monitor and evaluate capabilities, and measure risks and vulnerabilities via red teaming. In particular, we believe in a hybrid approach to test and evaluation which involves both automated evaluation and expert evaluations.
-

Red Teaming Methodology: Threat Modeling and Automated Risk Evaluation
By
–
For red teaming, the surface area of potential risks is extremely large. As a result, our approach involves extensive threat modeling, & automated evals to identify risks. Then our process enables capable experts to identify key risks, and report them back to model developers.
-
DEFCON Red Teams Leading AI Models for Security Vulnerabilities
By
–
At @DEFCON
, cybersecurity experts will be red teaming leading models from OpenAI, Anthropic, Google, and more. They will be using @scale_ai
’s platform to proactively identify vulnerabilities and report them to the model developers. -

Scale AI Launches LLM Test and Evaluation Platform
By
–
Today, alongside our collaboration with the @WhiteHouse and @DEFCON in an evaluation of the leading LLMs— @scale_ai is announcing the release of our Test and Evaluation platform and approach to enable for safe and scalable deployment of AI systems. Read thread for more
-
Authenticity and power: saying what you truly think
By
–
weird phenomenon— when you're young, you say what you think—you don't know any better! then as you get older, you start to say not what you think, but what others want you to say. eventually, you learn that true power IS saying exactly what you think, and revert to childhood.