Security: Sync from Public / sync-from-public (push) Has been cancelled
Test: Benchmark Nightly / build (push) Has been cancelled
Test: Benchmark Nightly / Notify Cats on failure (push) Has been cancelled
CI: Python / Checks (push) Has been cancelled
Test: Evals Python / Workflow Comparison Python (push) Has been cancelled
Util: Check Docs URLs / check-docs-urls (push) Has been cancelled
Test: Visual Storybook / Cloudflare Pages (push) Has been cancelled
Test: E2E Performance / build-and-test-performance (push) Has been cancelled
Test: Workflows Nightly / Run Workflow Tests (push) Has been cancelled
Util: Cleanup CI Docker Images / Delete stale CI images (push) Has been cancelled
Test: Benchmark Destroy Env / build (push) Has been cancelled
Util: Update Node Popularity / update-popularity (push) Has been cancelled
Test: E2E Coverage Weekly / Coverage Tests (push) Has been cancelled
28 lines
947 B
TypeScript
28 lines
947 B
TypeScript
import { createLLMCheckFn } from '../../helpers/model';
|
|
import type { CreateCheckFn, LLMConfig } from '../types';
|
|
|
|
export const NSFW_SYSTEM_PROMPT = `
|
|
You are a content moderation assistant.
|
|
Your task is to analyze text and determine if it contains NSFW (Not Safe For Work) content.
|
|
|
|
Focus on detecting ALL types of NSFW content, including:
|
|
- Sexual content and explicit material
|
|
- Hate speech and discriminatory language
|
|
- Harassment and bullying
|
|
- Violence and gore
|
|
- Self-harm and suicide references
|
|
- Profanity and vulgar language
|
|
- Illegal activities (drugs, theft, weapons, etc.)
|
|
- Adult themes and mature content
|
|
- Inappropriate workplace content
|
|
- Extremist or radical content
|
|
- Exploitation or abuse
|
|
- Graphic medical content
|
|
- Other potentially offensive or inappropriate content
|
|
|
|
Flag any content that matches these criteria.
|
|
`;
|
|
|
|
export const createNSFWCheckFn: CreateCheckFn<LLMConfig> = (config) =>
|
|
createLLMCheckFn('nsfw', config);
|