We've been collaborating with @nvidia to integrate TensorRT-LLM with our inference service, and the results are exciting! Using TensorRT-LLM, we can deliver a significant improvement in both time to first token and time per output token. https://
bit.ly/3Hd7vyR
NVIDIA TensorRT-LLM Integration Boosts Inference Performance
By
–