⚠️ CSAM, child abuse, violence, image generation
While there is current important reporting about the use of CSAM to train Grok, both original and generated content, I’d also like to issue the reminder that this is true for all large computational models where there is no way to verify all the content that goes into their making. People are, for some reason, just really adept at dismissing this from memory.
I could never bring myself to generate a generic photo using ”AI”, as I could never know if the software outputs the likeness of someone who has been abused, murdered or involved in an accident, child or adult. The potential further impact of this is horrifying to me.
2023: Stable Diffusion 1.5 Was Trained On Illegal Child Sexual Abuse Material, Stanford Study Says (Forbes)
https://www.forbes.com/sites/alexandralevine/2023/12/20/stable-diffusion-child-sexual-abuse-material-stanford-internet-observatory/
2023: Large AI Dataset has over 1,000 child abuse images, researchers find (Bloomberg)
https://www.bloomberg.com/news/articles/2023-12-20/large-ai-dataset-has-over-1-000-child-abuse-images-researchers-find
2025: Child sexual abuse and exploitation images identified in dataset used for developing AI moderation tools, analysis by Canadian child protection charity finds
https://www.protectchildren.ca/en/press-and-media/news-releases/2025/csam-nude-net
2025: How We Investigated the Epidemic of AI-Generated Child Sexual Abuse Material on the Internet
https://pulitzercenter.org/resource/how-we-investigated-epidemic-ai-generated-child-sexual-abuse-material-internet
Note: Once the computational model has been trained on this material, that’s it. You can’t take it out. You can’t ”unlearn” specific content from a computation, which is a further important distinction from deterministic software.
And remember that a huge amount of content for image generation ”training” is scraped from both digital and physical assets without any human intervention, so there are more ways for the content to enter than these already tainted datasets.
Call for filters? Automated dataset filtering is not a sufficient solution, and can introduce new biases:
2026: Dataset Filtering Provides Only Limited Protection Against CSAM Generation in Text-to-Image Models
https://cispa.de/en/cretu-dataset-filtering
What xAI’s Grok is allowing users to do, is double-down on this weakness and exploit it in full. The generated images are then treated as training data. All models have this weakness, but xAI does little do discourage it.
2026: Elon Musk’s xAI used child porn to train Grok models, lawsuit says
https://arstechnica.com/tech-policy/2026/08/elon-musks-xai-used-child-porn-to-train-grok-models-lawsuit-says/