this is the original RL with token-count penalty
RESEARCH
-

Reasoning in Latent States and Tokens for Transformers
By
–
What if you taught transformers to reason in both latent states and tokens? This Microsoft paper adds a self-supervised objective of predicting the next latent state to the standard next-token training, where a lightweight dynamic model
-
If Google had an enforceable patent on transformers
By
–
Imagine if Google had an enforceable patent on transformers.
-
Automating post-training with continual learning and OPSD-RL interpolation
By
–
on a technical level it feels like we're automating post-training we want to take a bunch of data and output a post-trained model – then iterate this process many times. continual learning by induction the ideal algorithm is an interpolation of OPSD and RL with a lot of
-
Mythos-level AI risks and open release preparation concerns
By
–
All Mythos-level models are likely to invite similar risks. Those risks will only be greater with the release of open Mythos-class AI coming in the next 6-12ish months (assuming China allows it) The lack of clarity over what risks concern the government may be slowing preparation
-
AI in 2040: Nearly Optimal Stack, Massive Current Inefficiency
By
–
AI in 2040 will not be built on the stack we use today. It will be much closer to optimal. The current stack exhibits 3-4 orders of magnitude data inefficiency and 4-5 orders of magnitude compute inefficiency. Nearly optimal AI is what
-

ArtiFixer: Generative 3D Scene Reconstruction from Text or Few Views
By
–
ArtiFixer doesn't just clean up 3D meshes—it is a robust generative engine. Even if you entirely drop the initial 3D rendering conditions, the model can rely on text prompts or just a few reference views to reconstruct the high-level structure of the scene and synthesize
-

ArtiFixer delivers sharp results, outperforming baselines by 1-3 dB PSNR
By
–
The results are incredibly sharp. When benchmarked on challenging datasets with sparse views like Mip-NeRF 360 and DL3DV, ArtiFixer handles highly degraded initial renderings effortlessly. It outperforms existing baselines (like GenFusion and 3DGUT) by a massive 1–3 dB PSNR
-

ArtiFixer uses DMD to make bidirectional video models 70x faster
By
–
Bidirectional video models provide great coherence but are computationally heavy. ArtiFixer solves this via Self-Forcing-style Distribution Matching Distillation (DMD). By distilling the bidirectional model into a causal auto-regressive one, ArtiFixer achieves up to a 70x
-

Opacity Mixing: Balancing Consistency and Hallucination in Video
By
–
How do you keep the generated video consistent with existing views without losing the ability to hallucinate new content? The answer is Opacity Mixing. Instead of starting from pure noise, ArtiFixer:
> Downscales the rendering's opacity map.
> Mixes Gaussian noise specifically