Analyze text content for potentially harmful material using the Azure Content Safety API. Returns severity scores for multiple harm categories including hate speech, sexual content, self-harm, and violence.
Arguments
- text
Character vector. The text(s) to analyze. Each text must be 10,000 characters or less.
- categories
Character vector. Categories to analyze. Must be a subset of
c("Hate", "Sexual", "SelfHarm", "Violence"). Default: all four categories.- output_type
Character. Severity level granularity. One of
"FourSeverityLevels"(returns 0, 2, 4, 6) or"EightSeverityLevels"(returns 0-7). Default:"FourSeverityLevels".- blocklists
Character vector of Content Safety blocklist names to apply.
- halt_on_blocklist
Logical. Whether the service should halt category analysis when blocklist content is found.
- endpoint
Character. Optional endpoint URL override. If NULL, uses the
AZURE_CONTENT_SAFETY_ENDPOINTenvironment variable.- api_key
Character. Optional API key override. If NULL, uses the
AZURE_CONTENT_SAFETY_KEYenvironment variable.- api_version
Character. API version to use. Default:
"2024-09-01".
Value
A tibble with columns:
- text
Character. The full input text.
- .input_idx
Integer. Position of the input in
text, useful for joining results back to caller data.- category
Character. The harm category: "Hate", "Sexual", "SelfHarm", or "Violence".
- severity
Integer. Severity score. Range depends on
output_type: 0-6 for FourSeverityLevels (values: 0, 2, 4, 6) or 0-7 for EightSeverityLevels.- label
Character. Human-readable severity label: "safe", "low", "medium", or "high".
- blocklist_matches
List. Blocklist matches returned by the service for the analyzed text.
- blocklist_hit
Logical.
TRUEwhen the service returned any blocklist match for the input text.- raw_response
List. Raw Content Safety response for the analyzed text.
Details
The Azure Content Safety API analyzes text for four types of harmful content:
Hate: Content that attacks or discriminates against individuals or groups based on protected attributes.
Sexual: Sexually explicit or adult content.
SelfHarm: Content that promotes or describes self-harm behaviors.
Violence: Content that describes or promotes violence.
Severity Labels:
safe (0-1): No harmful content detected.
low (2-3): Mildly concerning content.
medium (4-5): Moderately harmful content.
high (6-7): Severely harmful content.
The four-level scale uses the same labels at severities 0, 2, 4, and 6.
Authentication
You need an Azure Content Safety resource to use this function. Set up the endpoint and either an API key or a resource-scoped bearer-token provider:
Environment variables:
AZURE_CONTENT_SAFETY_ENDPOINTandAZURE_CONTENT_SAFETY_KEYHelper functions:
foundry_set_content_safety_endpoint()andfoundry_set_content_safety_key()Microsoft Entra ID:
foundry_set_token_provider()withscope = "resource"
Examples
if (FALSE) { # \dontrun{
# Requires an Azure Content Safety endpoint and credentials.
# Analyze a single text
foundry_moderate("This is a friendly message.")
# Analyze multiple texts
texts <- c(
"Hello, how are you today?",
"This is another message to check."
)
results <- foundry_moderate(texts)
# Filter for specific categories
foundry_moderate("Some text", categories = c("Hate", "Violence"))
# Use finer-grained severity levels
foundry_moderate("Some text", output_type = "EightSeverityLevels")
# Check results
library(dplyr)
results %>%
filter(severity > 0) %>%
arrange(desc(severity))
} # }