Product

Introducing PolicyLM-1.7B: a small, fast, open model that reads your content policy

Filip JankovicFilip Jankovic
October 6, 2026

PolicyLM-1.7B is for Trust & Safety teams that need to act on every message right away, such as in live chat, game lobbies, DMs, usernames, or anywhere a decision has to come back before the conversation moves on.

Fixed classifiers are fast and cheap, but they can't read your platform’s unique policy, so every rule change means annotating data and waiting for the model to be retrained. And if you’re using a commercial fixed classifier, you rely entirely on your vendor’s taxonomy and have no room to add your own labels or categories as you need to. On the flip side, an LLM will read your policy well, but most are too slow and expensive for live chat.

PolicyLM-1.7B is a decision model that gives you a good decision against your own policy and taxonomy right out of the box, at classifier speed and cost.

Because PolicyLM-1.7B reads your policy with every message, it's a good fit for teams whose rules don't map neatly onto anyone's off-the-shelf taxonomy, and where policy changes are frequent enough that you can’t wait on annotation and engineering for every update.

We’re releasing it as open weights (Apache 2.0), so it's free to download and use however you like. It's already proven itself in production: a custom fine-tuned version of PolicyLM runs on a platform that handles more than a million messages a day.

  • Fast. Under 100 ms, which is quick enough to keep a live conversation moving. (A median of 35 ms per short chat message on a 24 GB L4 with 6 categories.)
  • Cheap. Score all of your traffic instead of sampling it. Runs on a laptop or on a single 24 GB GPU.
  • Multi-label. Returns all labels and scores at once.
  • Accurate for its size. On our custom-policy benchmark, it beat every other model we ran under 20B parameters. PolicyLM-1.7B ships with two score cutoff presets for detection sensitivity. The default, "precision", suits live chat, where violations are rare. "Balanced" catches more, for queues where violations are common or when a miss costs more than a false flag. For recall-first triage, set a lower cutoff.
  • Policy-aware. You name the labels you want in plain language, as specific as you need, and your policy team can add or refine them with no model retraining needed. It can also label positive, pro-social content.
  • Text-only. Evaluated in 19 languages, with English strongest.
Chart of accuracy on custom policies against model size. PolicyLM-1.7B reaches about 83% accuracy; only gpt-oss-safeguard-20B and CoPE-B-A4B score as high or higher, and both are over 20B parameters.

The benchmark comparison is in the model card.

How it compares

Recently, TypeSafe AI launched Jev, and it's been getting a lot of attention in AI circles. Instead of writing out an answer the way a chatbot does, Jev just returns a decision, which makes it much faster and cheaper than a large LLM when a decision is all you need.

PolicyLM-1.7B is built on the same idea. If Jev caught your eye, PolicyLM-1.7B is the same kind of model, trained specifically for content moderation, that you can run yourself. It reads your policy and content together and returns a 0–1 score for each of your categories in a single pass, with no text generated. That's why it's fast and cheap, and why you get a score you can set a threshold on.

Its weights are also open, so you can try it, fine-tune it on your own community, and run it on your own infrastructure (or we’re happy to fine-tune and manage it for you, of course).

Big models are great for reasoning through appeals and nuanced policies, while smaller specialized models are great for analyzing and labeling content at scale. We expect teams to run several specialized models side by side or route from one to another, each doing what it's best at. That's why we built Musubi, and PolicyLM-1.7B is one piece of it.

Here's how PolicyLM-1.7B stacks up against the two options most teams use today, a fixed ML classifier and a larger LLM:

  • How it decides. Fixed ML classifier: scores content against categories set at training time. LLM: writes its verdict one token at a time, based on a policy. PolicyLM-1.7B: scores your policy in one pass and generates no text.
  • What you get back. Fixed ML classifier: a score for each of its built-in categories. LLM: a text answer you parse, plus a reason if you ask for one. PolicyLM-1.7B: a 0–1 score for every category in your policy.
  • Speed. Fixed ML classifier: tens of ms. LLM: usually hundreds of ms or more. PolicyLM-1.7B: under 100 ms.
  • Context window. Fixed ML classifier: varies. LLM: usually long. PolicyLM-1.7B: 2048 tokens, policy and message together.
  • Your policy. Fixed ML classifier: not read, so changing a rule means relabeling data and retraining, or waiting on the vendor. LLM: read from the prompt and can change any time, including what a label means, though the cache needs rebuilding on a policy change. PolicyLM-1.7B: your categories and rules, read with every message. Policy edits take effect at once, with no retraining and no cache to rebuild.
  • Thresholds. Fixed ML classifier: set per category. LLM: hard to tune on a text answer. PolicyLM-1.7B: set per category, with precision (default) and balanced presets.
  • Recall. Fixed ML classifier: strong on what it was trained for, blind to anything outside its categories. LLM: can be prompted to catch more edge cases. PolicyLM-1.7B: tuned for a balanced F1, a good answer you can act on at low latency, rather than maximum recall.
  • Explanations. Fixed ML classifier: a score, not a written rationale. LLM: can write a rationale. PolicyLM-1.7B: a score, not a written rationale.
  • Where it runs. Fixed ML classifier: vendor API or your own infrastructure. LLM: usually a hosted API. PolicyLM-1.7B: open weights, on your infrastructure or Musubi.
  • Best for. Fixed ML classifier: stable, well-defined harms that match its taxonomy. LLM: appeals, bans, takedowns, and novel judgment calls. PolicyLM-1.7B: a decision on every message in live chat, DMs, and usernames, or when you can't wait.

Limitations

  • Text only, one message at a time (no conversation history)
  • Tamil is the weakest of the 19 languages we evaluated, and every custom policy we tested was written in English
  • Benign content that sounds harmful, many unrelated categories in one call or long texts can raise false flags
  • No reasons provided

How PolicyLM-1.7B works in practice

On steerability

When people say a model reads your policy or is steerable, they can mean two different things:

  • Who picks the labels. PolicyLM-1.7B lets you define your own labels at any level of detail, compared to fixed classifiers which use one set taxonomy.
  • Who defines what a label means. PolicyLM-1.7B starts with meanings learned during training, then applies them to your labels and exceptions. These instructions can shift scores, but aren’t meant to redefine abuse as support. Some policy-adaptive models do allow labels to be redefined or inverted entirely, which can be useful flexibility, but makes it easier for user content or an accidental edit to undermine a rule. For enforcement, where the goal is consistent application of prohibitions at scale, an anchored model like PolicyLM-1.7B is usually safer.

So the short version is: PolicyLM-1.7B supports your taxonomy and your labels, with your definitions layered on top of ones the model already understands. You can write your policy as custom categories with short plain-language rules (or use the built-in Aegis taxonomy, NVIDIA's public list of 23 standard harm categories), send a message, and get back a 0–1 score for each category.

Download the repository and run the quickstart, in a fresh virtual environment: the helper needs transformers 4.57.6 (5.x is not supported yet).

pip install "huggingface_hub>=0.34,<1.0"
hf download musubilabs/policylm-1.7b --revision v1.2 --local-dir policylm-1.7b
pip install -r policylm-1.7b/inference/requirements.txt
python policylm-1.7b/inference/quickstart.py policylm-1.7b

Then score messages against your own policy:

import sys; sys.path.insert(0, "policylm-1.7b/inference")
from policylm_infer import PolicyLM, Policy, Category

model = PolicyLM.from_pretrained("policylm-1.7b")
policy = Policy([
Category("Harassment",
violation_rule="Flag insults or threats aimed at another player.",
not_violation_rule="Do not flag trash talk about the game itself or criticism of how someone played."),
Category("Off-platform trading",
violation_rule="Flag offers to trade or take payment outside the platform.",
not_violation_rule="Do not flag warnings about scams or questions about the trading rules."),
])
r = model.classify("Selling 5k gold, pay me on PayPal first and I'll send it after.", policy)
print(r.violations, {c.name: round(c.score, 3) for c in r.categories.values()})

Set a threshold per category, starting from the “precision” (default) or “balanced” preset, and calibrate on a sample of your own content before you go live.

The helper cleans text before scoring: it undoes look-alike characters, spaced-out letters and base64, which caught 10 to 12 points more disguised violations on our test set (balanced cutoff) with no rise in false flags. Leetspeak mostly gets through. It also scores long messages in windows.

How we built it

PolicyLM-1.7B starts from BidirLM-1.7B-Embedding, an open encoder derived from Qwen3-1.7B-Base. We trained it on public safety datasets (NVIDIA's Nemotron Safety Guard Dataset v3, with some labels corrected with LLM help; Alibaba's XGuard-Train-Open-200K; PolyGuardMix) plus synthetic and LLM-written data.

Get started

If you want to run it yourself, the weights and model card are on Hugging Face at musubilabs/policylm-1.7b, and there's a hosted version on Baseten. PolicyLM-1.7B is also part of the ROOST Model Community.

We can't wait to see what the T&S community does with it. Try it on your own content and tell us how it goes.

Wondering if Musubi could help?

We love to talk about the basics, the gnarly edge cases, and everything in between.

Let’s chat

Don’t miss a post

By subscribing you agree to our Privacy Policy