Our algorithm trains LLM classification systems to block harmful inputs and outputs based on a “constitution” of harmful and harmless categories of information.
LLM Classification Systems Trained on Constitutional Harm Prevention
By
–

By
–

Our algorithm trains LLM classification systems to block harmful inputs and outputs based on a “constitution” of harmful and harmless categories of information.