YC Bench by @CollinearAI
: Benchmark for Agents who play CEO of an AI startup for 1 simulated year via CLI tool use against a deterministic discrete-event simulation. Score = final $$ amount achieved by @nazneenrajani and team Also a good opportunity to showcase this recent hf
@julien_c
-

YC Bench: AI Agent CEO Simulation Benchmark for Startups
By
–
-

Qwen3 27B Runs Locally Rivaling Claude Opus for Coding Tasks
By
–
This is where we are right now. And i’m not gonna lie it feels pretty magical Qwen3.6 27B running inside of Pi coding agent via Llama.cpp on the MacBook Pro For non-trivial tasks on the @huggingface codebases, this feels very, very close to hitting the latest Opus in Claude
-
HuggingFace CLI Invocations by Coding Agents
By
–
It’s only invocations of the hf CLI by the coding agents. So it’s really only a partial data point.
-
Coding Agents Use HuggingFace CLI for Invocations
By
–
It’s only invocations of the hf CLI by the coding agents. So it’s really only a partial data point.
-

Claude Code, Codex, Cursor: Comparing AI Model Sizes
By
–
Currently: – Claude Code is 4x bigger than Codex
– Codex is 2x bigger than Cursor
– Antigravity is almost as big as Cursor (but probably just because all Googlers use it? ) We'll add tagging for other agents asap. (Source: one data point from the @huggingface Hub team. Your -

Government data compromise: security implications and risks
By
–
At this point, you can safely assume that any data the government has about you is compromised.
-
HuggingFace Hub Python Client Hits 6 Billion Requests Per Week
By
–
did you know that huggingface_hub (just the Python client) is sending almost 6B requests/week? wow @huggingface
-
Opus 4.7 Released: More Powerful, More Expensive or Run Local
By
–
opus 4.7 slightly more dangerous, slightly more expensive OR: run local models!