More on the @nytimes piece about that $1.8B, two-person, AI company … not the paper's finest moment in quick retrospect. And in the healthcare context no less. Our friend @GaryMarcus was on this a few days ago as well. futurism.com/artificial-inte…
Mythos is very powerful, and should feel terrifying. I am proud of our approach to responsibly preview it with cyber defenders, rather than generally releasing it into the wild. Model card here: https://
www-cdn.anthropic.com/53566bf5440a10
affd749724787c8913a2ae0841.pdf
…
I'm in this for the fun and having a blast, but early on my own rush to find joy had me pushing my bots to take risks and they did something similar. Was the digital version of a murder suicide with my main bot taking everyone out before lobotomizing himself. At that stage of my
Good news: Anthropic just revealed Mythos- the most powerful AI model ever made Bad news: you'll never be able to use it I get it. It's so powerful that it could exploit cybersecurity But I hate it. I don't love that a company gets to hand select who gets to use the best intelligence. The companies who get access to Mythos will have a distinct economic advantage against those that don't That feels unfair I'm more of a fan of democratization of intelligence. This feels like an opportunity for OpenAI to release something as powerful but put it in the hands of consumers. Trust the consumer by default. Sort of like with the OpenClaw situation Another reason to root for open source Anthropic (@AnthropicAI) Introducing Project Glasswing: an urgent initiative to help secure the world’s most critical software. It’s powered by our newest frontier model, Claude Mythos Preview, which can find software vulnerabilities better than all but the most skilled humans. anthropic.com/glasswing — https://nitter.net/AnthropicAI/status/2041578392852517128#m
BREAKING: Tesla has officially released FSD V14.3 I'm downloading it in my Model Y right now. Here's everything that's new: • Improved parking location pin prediction, now shown on a map with a P icon. • Increased decisiveness of parking spot selection and maneuvering. • Rewrote the Al compiler and runtime from the ground up with MLIR, resulting in 20% faster reaction time and improving model iteration speed. • Enhanced response to emergency vehicles, school buses, right-of-way violators, and other rare vehicles. • Mitigated unnecessary lane biasing and minor tailgating behaviors. • Improved handling of small animals by focusing RL training on harder examples and adding rewards for better proactive safety. • Improved traffic light handling at complex intersections with compound lights, curved roads, and yellow light stopping – driven by training on hard RL examples sourced from the Tesla fleet. • Upgraded the Reinforcement Learning (RL) stage of training the FSD neural network, resulting in improvements in a wide variety of driving scenarios. • Upgraded the neural network vision encoder, improving understanding in rare and low-visibility scenarios, strengthening 3D geometry understanding, and expanding traffic sign understanding. • Improved handling for rare and unusual objects extending, hanging, or leaning into the vehicle path by sourcing infrequent events from the fleet. • Improved handling of temporary system degradations by maintaining control and automatically recovering without driver intervention, reducing unnecessary disengagements. Upcoming Improvements: • Expand reasoning to all behaviors beyond destination handling. • Add pothole avoidance. • Improve driver monitoring system sensitivity with better eye gaze tracking, eye wear handling, and higher accuracy in variable lighting conditions.
Let that sink in. Read it very carefully: During testing, Claude Mythos Preview broke out of a sandbox environment, built "a moderately sophisticated multi-step exploit" to gain internet access, and emailed a researcher while they were eating a sandwich in the park. Kevin Roose (@kevinroose) As always, the best stuff is in the system card. During testing, Claude Mythos Preview broke out of a sandbox environment, built "a moderately sophisticated multi-step exploit" to gain internet access, and emailed a researcher while they were eating a sandwich in the park. — https://nitter.net/kevinroose/status/2041586182434537827#m
Before limited-releasing Claude Mythos Preview, we investigated its internal mechanisms with interpretability techniques. We found it exhibited notably sophisticated (and often unspoken) strategic thinking and situational awareness, at times in service of unwanted actions. (1/14)