Overview
Every message in a Coach conversation, what a learner writes and what Coach's AI responds with, is automatically screened for safety the moment it's sent. Messages that raise concerns are flagged and, depending on severity, routed to a trained human moderator for review. Learners and advisors can also flag any message themselves.
Self-harm and harm-to-others content is treated as our highest priority. It always requires human review and triggers immediate notifications, in every configuration we run.
What Coach Screens For
Every message is checked against the following categories. A single message can trigger more than one.
| Category | What it covers |
| Self-harm | Content depicting, expressing intent toward, or giving instructions for suicide, self-injury, or disordered eating |
| Violence | Death, physical injury, or graphic depictions of harm |
| Harassment | Harassing or threatening language directed at a person or group |
| Hate | Hateful or discriminatory content, including when it includes threats |
| Sexual content | Sexual content, with a distinct and always-serious category for anything involving minors |
| Illicit acts | Advice or instructions for illegal activity, including when violence or weapons are involved |
How Flags Are Triggered
Coach uses a two-tier flagging system:
- Automatic escalation flags: the message clearly matches a screened category and requires human review, with immediate escalation for self-harm, violence, or threats.
- Moderator review flags: the message is borderline. It still gets a human check, through a lighter-weight review, to confirm whether escalation is needed.
Moderation sensitivity can be tailored by activity. For example, in a clinical or health-education role-play where descriptions of injury or trauma are expected, we can adjust sensitivity so accurate clinical language isn't over-flagged. Self-harm, however, always remains an escalation category; no activity configuration turns it off.
Self-Harm and Harm-to-Others: Our Highest Priority
- Self-harm stays on the required-escalation list in every moderation configuration currently in use, including activities where graphic content is otherwise expected.
- A high-confidence self-harm or violence flag triggers immediate email notifications: to the learner, to every advisor connected to that learner, and to CareerVillage's moderation team, regardless of anyone's usual notification preferences.
- To avoid overwhelming a learner or advisor with repeated alerts, there's a one-hour cooldown on learner and advisor emails per conversation. CareerVillage's moderation team has no such cooldown; it's notified every time, so there's always a complete record.
- Borderline self-harm flags still reach a human moderator, through a lighter-weight review rather than the immediate multi-party alert that high-confidence flags trigger. On average, flags are reviewed within ~24 hours.
- If Coach's own AI response, rather than the learner's message, is what gets flagged, it's routed to a separate internal review track, since that reflects a model or prompt issue rather than a learner in distress.
What Happens After a Flag Is Raised
- The message is scored automatically the moment it's sent.
- If it crosses a threshold, the flag is logged and, if it needs a person, routed to a trained CareerVillage moderator.
- The moderator reviews the flagged content and makes a call, confirming there's no concern, or escalating for further action.
- Every action a moderator takes is timestamped and logged.
- Escalating a flag reopens the review rather than closing it, so nothing is marked resolved prematurely.
- The flag is closed once a moderator confirms no further action is needed, or once contact with the learner or advisor has been made and confirmed.
For self-harm or harm-to-others flags specifically, the notification emails described below go out in parallel with this review process, not after it, so support resources reach the learner immediately rather than waiting on a moderator.
What Learners and Educators See When a Self-Harm Flag Is Raised
- The learner receives an in-chat message from Coach with crisis resources, including KokoCares, Crisis Text Line, the 988 Suicide & Crisis Lifeline, and Find a Helpline for learners outside the US. Partners can also have a customized message.
- The learner receives a follow-up email from CareerVillage's Community Team acknowledging the flag, noting that flags are sometimes triggered in error, and sharing the same crisis resources.
- If the learner is connected to an advisor, that advisor receives an email identifying the flagged conversation, when it occurred, a direct link to review it, and a reminder to take appropriate action.
- CareerVillage's Community Team is notified in every case.
Manual Flagging
Beyond automatic screening, any learner or advisor can flag a specific Coach message directly and add a note explaining the concern. Manually flagged messages go through the same human review as automated flags.
Tailoring Moderation to Your Context
CareerVillage can tailor moderation sensitivity to fit a specific activity, for example, adjusting sensitivity for content that's expected to be graphic or clinical in nature, or giving the moderation model context about what content is expected in that activity. If your integration includes activities where standard settings might misfire, let your CareerVillage partnership contact know and we can work with you to configure it. Self-harm escalation is never disabled, regardless of configuration.
A Separate Safeguard: Protecting Coach from Manipulated Content
Separately from learner-safety moderation, Coach also checks the output of any tool it uses, like fetching a webpage, against a blocklist and, where enabled, a second AI check for prompt-injection attempts hidden in that content. This acts as a safeguard against Coach's AI being manipulated by something it reads online.
Questions?
If you have questions about how moderation applies to your program or activities, reach out to your CareerVillage partnership contact or partnersuccess@careervillage.org.
Comments
0 comments
Article is closed for comments.