> ## Documentation Index
> Fetch the complete documentation index at: https://learn.social.plus/llms.txt
> Use this file to discover all available pages before exploring further.

# AI Content Moderation

> Leverage AI-powered moderation tools to create safer communities with automated content filtering and intelligent threat detection

Create safer online communities with intelligent, automated content moderation. social.plus leverages advanced AI to scan and filter inappropriate content across text, images, and video, ensuring community standards are maintained without constant manual oversight.

<CardGroup cols={2}>
  <Card title="Pre-Moderation" icon="shield-check" href="#ai-pre-moderation">
    Block inappropriate content before it's published with proactive AI scanning
  </Card>

  <Card title="Post-Moderation" icon="eye" href="#ai-post-moderation">
    Monitor and review published content with intelligent flagging and automated actions
  </Card>
</CardGroup>

## Overview

social.plus offers two complementary AI moderation approaches:

<AccordionGroup>
  <Accordion title="Pre-Moderation" icon="shield-check">
    **Proactive Content Filtering**

    * Content is scanned before publication
    * AI generates confidence scores for detected violations
    * Content blocked if confidence exceeds configured threshold
    * User must modify content to proceed with posting
  </Accordion>

  <Accordion title="Post-Moderation" icon="eye">
    **Reactive Content Review**

    * Content is scanned after publication
    * Uses `flagConfidence` and `blockConfidence` thresholds
    * Automatically flags content for review or removes violations
    * Maintains community safety without blocking legitimate content
  </Accordion>
</AccordionGroup>

## Getting Started

<Steps>
  <Step title="Enable AI Moderation">
    Contact our [support team](mailto:support@social.plus) to enable AI content moderation for your application.
  </Step>

  <Step title="Configure Settings">
    Set up confidence levels and moderation categories through the social.plus Console.
  </Step>

  <Step title="Test & Monitor">
    Test with sample content and monitor moderation effectiveness through analytics.
  </Step>
</Steps>

## AI Pre-Moderation

Prevent inappropriate content from reaching your community with proactive AI scanning. Pre-moderation ensures all content meets your standards before publication.

<Info>
  **Current Availability**: Pre-moderation is currently available for image content, with text and video support coming soon.
</Info>

### Image Content Detection

Our AI pre-moderation scans all uploaded images for inappropriate content across four key categories:

<AccordionGroup>
  <Accordion title="Content Categories">
    * **Nudity**: Detection of explicit or inappropriate nudity
    * **Suggestive Content**: Sexually suggestive or provocative imagery
    * **Violence**: Violent or graphic content detection
    * **Disturbing Content**: Content that may be psychologically disturbing
  </Accordion>
</AccordionGroup>

### Configuration

<Steps>
  <Step title="Enable Image Moderation">
    Navigate to **Moderation > Image Moderation** in your social.plus Console and toggle "Enable image moderation" to **ON**.
  </Step>

  <Step title="Set Confidence Levels">
    Configure confidence thresholds for each category based on your community standards.
  </Step>

  <Step title="Test Configuration">
    Upload test images to verify your confidence settings work as expected.
  </Step>
</Steps>

### Understanding Confidence Levels

<Warning>
  **Important**: Confidence levels significantly impact moderation accuracy. Default settings may produce false positives.
</Warning>

Confidence levels represent the AI's certainty in detecting specific content types:

* **Low Confidence (0-30)**: High sensitivity, may block legitimate content
* **Medium Confidence (40-70)**: Balanced approach for most communities
* **High Confidence (80-100)**: Conservative filtering, may miss some violations

<Tip>
  **Recommendation**: Start with medium confidence levels (40-60) and adjust based on your community's needs and false positive rates.
</Tip>

## AI Post-Moderation

Monitor and moderate published content with intelligent detection and automated response workflows. Post-moderation provides comprehensive scanning across all content types while maintaining user experience.

All AI post-moderation results are surfaced through the **Moderation Feed** in the social.plus Console, located under **Moderation > Moderation feed**.

<CardGroup cols={3}>
  <Card title="Text Moderation" icon="message" href="#text-content-detection">
    Detect inappropriate language, hate speech, and harmful text content
  </Card>

  <Card title="Image & Video" icon="image" href="#multimedia-content-detection">
    Scan visual content for policy violations and harmful imagery
  </Card>

  <Card title="User Profile Moderation" icon="user-shield" href="/analytics-and-moderation/console/ai-user-profile-moderation">
    AI moderation for display names, avatars, and descriptions — see dedicated page
  </Card>
</CardGroup>

### Moderation Feed

The Moderation Feed is the central hub for reviewing AI-flagged content. It is organized into two main workflow tabs:

<Tabs>
  <Tab title="To Review">
    The **To review** tab displays all content that requires moderator attention, organized into sub-tabs:

    * **Posts and comments** — Flagged posts and comments from communities and user timelines
    * **Messages** — Flagged messages from channels and direct conversations
    * **Users** — Flagged user profiles — see [AI User Profile Moderation](/analytics-and-moderation/console/ai-user-profile-moderation)

    Each flagged item displays:

    * The AI moderation label and detected categories (e.g., "AI Mod: Harassment or bullying")
    * PII detection results when applicable (e.g., URLs, person types)
    * The number of user reports (e.g., "1 user", "4 users")
    * The last flagged timestamp
    * Available moderation actions

    **Available actions for posts and messages:**

    * **Delete post / Delete message** — Remove the content
    * **Clear flag** — Dismiss the flag and approve the content
  </Tab>

  <Tab title="Reviewed">
    The **Reviewed** tab shows content that has already been moderated, organized into the same three sub-tabs (Posts and comments, Messages, Users).

    Each reviewed item shows the action taken and who performed it:

    * **Flag cleared by \[moderator]** — Flag was dismissed
    * **Deleted by \[moderator]** — Content was removed
    * **Reviewed by \[moderator]** — Content was reviewed without a specific action

    The Reviewed tab includes an additional **Select moderator** filter to view actions taken by a specific team member.
  </Tab>
</Tabs>

<Tip>
  Use the filter dropdowns (All feeds / All Channels, Select creator / Select sender) to narrow down the moderation queue by community, channel, or content creator.
</Tip>

### Content Coverage

<AccordionGroup>
  <Accordion title="Posts and Comments">
    * Text, images, videos, clips, files, and livestream content
    * Posts across all communities and user timelines
    * Comments and reply chains on posts
    * Filter by specific community feed or content creator
  </Accordion>

  <Accordion title="Messages">
    * Text, image, video, audio, and file messages
    * Messages from group channels, direct messages, and live chat
    * Filter by specific channel or message sender
  </Accordion>
</AccordionGroup>

### Text Content Detection

The AI text moderation identifies and handles various types of inappropriate text content:

<AccordionGroup>
  <Accordion title="Detection Categories" icon="text">
    * **Harassment or Bullying**: Targeted abuse, intimidation, or bullying behavior
    * **Sexual Content or Nudity**: Adult content and explicit sexual references
    * **Violence or Threatening Content**: Violent threats, graphic descriptions, or dangerous activities
    * **Hate**: Hate speech targeting protected groups
    * **Fraudulent Intent and Scam Promotion**: Scam tactics, phishing, and deceptive content
    * **Self Harm or Suicide**: Content related to self-harm or suicidal ideation
  </Accordion>

  <Accordion title="PII Detection" icon="fingerprint">
    In addition to content policy categories, the AI performs **Personally Identifiable Information (PII) detection** to flag sensitive data:

    * **URL** — Links and web addresses embedded in content
    * **PersonType** — References to specific person types or identities
  </Accordion>
</AccordionGroup>

<Info>
  Content that passes all AI checks displays **"AI Mod: Passed"** in the moderation feed. Content with detected violations shows the specific category labels.
</Info>

### Multimedia Content Detection

<Warning>
  **Comprehensive Scanning**: Our AI analyzes both static images and video content frame-by-frame for maximum protection.
</Warning>

Advanced visual content analysis covers extensive categories:

<AccordionGroup>
  <Accordion title="Adult Content" icon="eye-slash">
    * Adult Toys, Explicit Nudity, Graphic Nudity
    * Sexual Activity, Sexual Situations, Suggestive Content
    * Female Swimwear or Underwear, Swimwear or Underwear
    * Non-Explicit Nudity of Intimate Parts and Kissing
    * Partial Nudity, Illustrated Explicit Nudity, Revealing Clothes
  </Accordion>

  <Accordion title="Violence & Harmful Content" icon="shield-exclamation">
    * Violence, Graphic Violence, Gore, Physical Violence
    * Weapons, Weapon Violence, Explosions
    * Self Injury, Hanging, Corpses
    * Emaciated Bodies, Visually Disturbing Content
  </Accordion>

  <Accordion title="Substance-Related Content" icon="flask">
    * Alcohol, Alcoholic Beverages, Drinking
    * Drugs, Drug Products, Drug Use, Drug Paraphernalia
    * Pills, Smoking, Tobacco, Tobacco Products
  </Accordion>

  <Accordion title="Extremist & Hate Content" icon="ban">
    * Hate, Extremist, Nazi Party, White Supremacy
    * Hate Symbols, Rude Gestures, Middle Finger
  </Accordion>

  <Accordion title="Other Restricted Content" icon="triangle-exclamation">
    * Gambling, Air Crash, Disasters
    * Bare-chested Male (context-dependent)
    * Other contextually inappropriate content
  </Accordion>
</AccordionGroup>

<Info>
  **User Profile Moderation**: AI moderation for user profiles (display names, avatars, descriptions) is covered in a dedicated page. See [AI User Profile Moderation](/analytics-and-moderation/console/ai-user-profile-moderation) for setup, admin reset workflows, blocklist configuration, and moderation feed details.
</Info>

### Understanding Confidence Scores

<AccordionGroup>
  <Accordion title="Confidence Thresholds" icon="sliders">
    **Flag Confidence** (Default: 40)

    * Content scoring above this level gets flagged for review
    * Lower values = more content flagged (higher sensitivity)
    * Recommended range: 30-60 depending on community standards

    **Block Confidence** (Default: 80)

    * Content scoring above this level gets automatically removed
    * Higher values = fewer false positives
    * Recommended range: 70-90 for balanced protection
  </Accordion>

  <Accordion title="Score Ranges" icon="chart-bar">
    * **0-39**: Content passes moderation (approved)
    * **40-79**: Content flagged for human review
    * **80-100**: Content automatically blocked/removed

    *Note: These ranges use default thresholds and can be customized*
  </Accordion>
</AccordionGroup>

<Info>
  **Default Configuration**: All categories start with `flagConfidence: 40` and `blockConfidence: 80`. Monitor your community's content patterns and adjust these values to optimize for your specific needs.
</Info>

### Configuration Parameters

<AccordionGroup>
  <Accordion title="Parameter Reference" icon="gear">
    | Parameter         | Type   | Description                            |
    | ----------------- | ------ | -------------------------------------- |
    | `category`        | String | Name of the moderation category        |
    | `flagConfidence`  | Number | Threshold for flagging content (0-100) |
    | `blockConfidence` | Number | Threshold for blocking content (0-100) |
    | `moderationType`  | String | Type of content: "text" or "media"     |
  </Accordion>
</AccordionGroup>

### API Configuration

<Tabs>
  <Tab title="Regional Endpoints">
    Select the appropriate API endpoint for your region to ensure optimal performance:

    | Region            | API Endpoint                  |
    | ----------------- | ----------------------------- |
    | **Europe**        | `https://api-eu.social.plus/` |
    | **Singapore**     | `https://api-sg.social.plus/` |
    | **United States** | `https://api-us.social.plus/` |
  </Tab>

  <Tab title="Configuration APIs">
    <CodeGroup>
      ```json Retrieve confidence level theme={null}

      GET /v1/content-moderation/confidences

      Response:
      {
        "categories": [
          {
            "category": "explicit_content",
            "flagConfidence": 40,
            "blockConfidence": 80,
            "moderationType": "media"
          }
        ]
      }
      ```

      ```json Update Moderation Confidence theme={null}
      PUT /v1/content-moderation/confidences

      Request:
      {
        "category": "explicit_content",
        "flagConfidence": 50,
        "blockConfidence": 85
      }
      ```
    </CodeGroup>
  </Tab>
</Tabs>

## API Reference

<Note>
  For detailed administration workflows, see the [Moderation Overview](/analytics-and-moderation/console/moderation/overview) and analytics export documentation.
</Note>

## Best Practices

<AccordionGroup>
  <Accordion title="Configuration Strategy" icon="gear">
    * **Start Conservative**: Begin with moderate confidence levels and adjust based on results
    * **Monitor Performance**: Track false positive and false negative rates
    * **Community-Specific**: Tailor settings to your community's content standards
    * **Regular Review**: Periodically review and update thresholds as your community evolves
  </Accordion>

  <Accordion title="Human Oversight" icon="users">
    * **Review Queue Management**: Ensure consistent review of flagged content
    * **Moderator Training**: Train team on community standards and edge cases
    * **Appeal Process**: Provide clear paths for users to contest moderation decisions
    * **Transparency**: Communicate moderation policies clearly to users
  </Accordion>

  <Accordion title="Performance Optimization" icon="gauge">
    * **Batch Processing**: Handle high-volume content efficiently
    * **Regional APIs**: Use geographically appropriate endpoints
    * **Webhook Integration**: Implement real-time event handling for flagged content
    * **Monitoring**: Set up alerts for unusual moderation patterns
  </Accordion>
</AccordionGroup>
