Guides

How to Choose and Route Open Weights Content Moderation Models

Musubi
October 9, 2026

To choose an open-source content moderation model, test candidates on your own policy and data for accuracy, cost at your scale, speed, and how faithfully they follow your rules. No single model is best for every task, so many teams route content through a small, fast model first and escalate only uncertain or high-risk cases to a larger one. Start simple, and add routing only when cost or latency requires it.

Together with ROOST, we've published Choosing and Routing Open Moderation Models, a free practitioner report on selecting open moderation models and combining them into workflows that fit your policies and operating constraints. Below are the core principles, and the full guide has the details.

How to choose an open moderation model

Four questions usually decide whether a model is viable: how accurate it is on your policy, what it costs at your scale, how fast it is, and how easily you can steer it as your policy changes.

  • Match the accuracy metric to the model's role. Favor recall for a first pass, since a later stage can resolve false positives. Favor precision for a second layer, where false positives trigger enforcement or human review. Aim for balanced F1 when a model decides on its own.
  • Test on your own policy and traffic. Public benchmarks help narrow the field, but a golden set of examples your team has reviewed and agreed on is what tells you whether a model works for you.
  • Set the latency budget by surface. Live chat, DMs, and usernames need decisions within tens to low hundreds of milliseconds, while uploads and profiles can usually tolerate a second or more. Measure end-to-end system time, not just model time.
  • Spend model budget where it matters most. Where you spend matters more than which model you choose, and automatically approving clean content is often the largest operational saving.
  • Expect to tune policy language for each model. Steerable models follow custom policies with varying fidelity, and the same policy can behave very differently across model sizes and versions.
  • Test each language and adversarial pressure separately. "Multilingual" on a model card doesn't mean your policy performs equally well in every locale, and any model that reads policy should be tested against fake policy text injected into content.

When to route between models

Routing works well when most content is obviously safe, when hard subcategories benefit from specialist policies, or when your budget can't cover a large model on all traffic. A single model is the better fit when decisions must return in one pass, as in live chat, when volume is low, or when your team doesn't yet have a golden set to tune handoffs safely.

When you do route:

  • Tune the first pass for recall. Anything it clears is never reviewed again.
  • Let labels decide where content goes, and scores decide whether it goes anywhere. Allow content below a low threshold, act automatically above a high one, and escalate what falls in between.
  • Calibrate thresholds by severity, language, position in the stack, and model. A 0.7 from one model isn't equivalent to a 0.7 from another.
  • Log every score and version every policy so you can test new thresholds without rerunning traffic and revert quickly when something breaks.

Start with the simplest workflow that meets your cost and latency needs, and add complexity only when you have to. The largest model isn't needed everywhere, and small, open-source models are often sufficient.

Read the full guide

The full guide, by Musubi’s Alice Hunsberger and Juliet Jonak, and ROOST’s Jeremie Ponak, goes deeper on each of these principles. It includes findings from Musubi and ROOST testing, reference routing architectures, example stacks, and a summary of models in the ROOST Model Community.

Read the guide →

Wondering if Musubi could help?

We love to talk about the basics, the gnarly edge cases, and everything in between.

Let’s chat

Don’t miss a post

By subscribing you agree to our Privacy Policy