Guides

A Field Guide to the Trust & Safety Vendor Landscape

Alice HunsbergerAlice Hunsberger
August 28, 2026

Trust & safety vendors tend to get lumped into one category, but there are at least five distinct kinds: automated content decisioning, platforms for moderation operations, identity verification/age assurance, fraud analytics and automation, and managed services that supply reviewers. A good evaluation starts by knowing which kind you're looking at. A labeling API and an outsourced review team solve different problems and have different pros and cons. This guide covers what each one is, where it fits, and how to match it to the bottleneck you have.

Several of these companies span more than one category (including us); we've listed each where they belong. You should decide what makes sense for you.

Automated Content Decisioning

You send content, you get back labels or scores, and then you decide what to do with that label next. APIs handle volume and latency well, they're much cheaper than human moderation, and integration is easy, which is why most platforms start here. There are three types of these:

Fixed Classifiers

The classic version of a classifier API has a fixed taxonomy, which is a real structural limit. The categories were trained on the vendor's definition of harm, and those labels might not match up with what's important to you, your users, or your platform. When your policy and the label set disagree, you have to layer on human review, which can get expensive.

Companies offering traditional fixed classifiers include Hive, Sightengine, Amazon Rekognition, Azure AI Content Safety, the OpenAI moderation endpoint, Perspective API (which is sunsetting), and Community Sift (also sunsetting).

There are also specialty classifiers in the child safety space to surface CSAM (child sexual abuse material). These include Google's Child Safety Toolkit, Microsoft PhotoDNA, IWF, Safer, and Resolver. (Musubi natively integrates with the first three).

Policy-Aware Classifiers

A newer variant of a moderation API is policy-aware classifiers, which are small models conditioned on your written policy rather than a fixed category list. Still API-shaped, still cheap and fast, but the taxonomy problem moves from "unfixable" to "editable."

Examples of policy-aware classifier models are Musubi's PolicyLM-1B model, Zentropi's CoPe models, and Mistral's Shieldstral model.

LLM Moderation

LLM moderation is the third tier, and it's automated content moderation too, just not a cheap classifier. Instead of returning a label, the model reads your written policy against the content and returns a decision along with its reasoning and a confidence score. It costs more per item than a classifier, which is why teams tend to reserve it for nuanced policies and hard cases, but it costs far less per policy change: edit the policy and the behavior changes with it, no retraining cycle. The reasoning is the part people underestimate. When a decision gets appealed or a regulator asks why, a written rationale is an answer. A label with a score isn't.

Companies offering LLM moderation include Musubi, Moonbounce, and SafetyKit. Some teams DIY this first if their needs are simple, using GPT, Claude, Gemini, or open-source safety models like the ones in Roost's Model Community.

Whichever type you use, it's worth asking what happens to content the model isn't sure about, or doesn't label correctly, and what you want escalated for human review and oversight.

Resources:

Operations Platforms

This is what most people picture when they say T&S tooling: the console where moderators work. Features will include report queues, case management, child safety review, and appeals queues. If the API answers "what is this content," the platform answers "what happens next, and who decided."

Platform vendors focus on efficient human review. One question to ask is how much human review do you actually want to be doing? A platform organized around the queue assumes the queue is the work. Some operations platforms do include automation as an add-on, but it's not their central focus.

Companies offering operations platforms include Coop (open source; Musubi offers a hosted version unsurprisingly called Musubi Coop), Tremau (integrated with Musubi), Checkstep, and Cinder.

Identity Verification/Age Assurance

A different job entirely: not "what is this content" but "who is this person." This includes document checks, age assurance, liveness, and sometimes full KYC. These vendors show up in T&S evaluations because regulators increasingly ask platforms to know things about their users, and because fake accounts are upstream of most abuse.

If your problem is content or behavior, an identity vendor won't solve it. If your problem is a regulator asking how you keep minors out, you probably need an identity or age check vendor.

Examples of identity verification vendors include Persona and Yoti.

Fraud Analytics and Automation

Fraud analytics is risk scoring on accounts and behavior rather than individual pieces of content. This includes looking at device signals, network patterns, velocity, and connections between users. This is how you catch the actor, not just individual pieces of content.

Fraud models are either created with shared data or your own data. Shared data means faster cold starts and learning from someone else's fraud patterns. Your own data means the system reflects how your team rules on cases and your data stays yours.

Examples of fraud vendors are Musubi, Sift, Sardine, SafetyKit, and Incognia.

Managed Review Services (BPOs)

BPOs supply trained moderators at whatever volume you contract for. For a decade this was the default answer to "our queue is crushing us." The tradeoffs are operational. Training and quality assurance for a large external team is hard work, and accuracy is hard to audit from the outside. Contracts price on headcount and hours and can add up quickly, and the incentive to automate sits with you. However, many BPOs can bring decades of expertise from previous partnerships, including insights on policy, language and culture, and best practices.

Examples of BPOs in the moderation space include TaskUs, Concentrix, Teleperformance, Alorica, Modsquad, and Highspring.

Resource: The Future of T&S is Better Collaboration from Musubi + Alorica

How to decide what you need next

If you're not satisfied with your current moderation system, but don't know where to start, figure out what your main bottleneck is.

  • If you're overwhelmed by content or your human review is getting too expensive, look into automation.
  • If you need somewhere to organize appeals and child safety reports, look into a moderation console.
  • If you have fake accounts that are multiplying daily, you need fraud analytics.
  • If a regulator is asking you to verify age on your platform, you need identity verification or age assurance.
  • If you need language expertise or 24/7 moderation that your internal team can't cover, look into a BPO.
  • If you have a bunch of fractured tools and they're not speaking to each other, look for a vendor that covers multiple platforms.

Whichever you choose next, it's important to understand not only what services the vendor offers, but what else they do, how they integrate with other platforms, what expertise and support they bring to the table, and how they play well with others. A good vendor should be transparent, collaborative, and focused on solving your problems first and foremost, not just making an easy sale.

Wondering if Musubi could help?

We love to talk about the basics, the gnarly edge cases, and everything in between.

Let’s chat

Don’t miss a post

By subscribing you agree to our Privacy Policy