Can AI training be fixed by keeping data on the chip longer? Researchers from MIT, Princeton, Together AI, and Meta introduce CODA — a new way to rewrite Transformer building blocks as GEMM-plus-epilogue programs. Instead of moving large intermediate tensors back and forth to
Researchers propose CODA to keep data on chip longer for AI training
By
–
