🚨 @NVIDIA just dropped LocateAnything, making object detection ~10x faster by fixing one core bottleneck:
— Charly Wargnier (@DataChaz) 17 juin 2026
How the model writes coordinates.
Standard AI models do visual grounding the slow way.
They predict coordinates piece by piece: Token 1, Token 2, Token 3, Token 4 etc.… pic.twitter.com/GAyr0taCbX
@NVIDIA just dropped LocateAnything, making object detection ~10x faster by fixing one core bottleneck: How the model writes coordinates. Standard AI models do visual grounding the slow way. They predict coordinates piece by piece: Token 1, Token 2, Token 3, Token 4 etc.