Scribd, Inc. Classifies More Than 400 Million Documents with Gemini Batch Inference on Gemini Enterprise
Google, Thursday, September 24th, 2026
Scribd used Gemini's PDF understanding and Gemini Enterprise batch inference to classify over 400 million user-uploaded documents.
Scribd used Gemini Enterprise's native PDF understanding and batch prediction to run trust-and-safety classification across its entire user-generated content corpus of more than 400 million documents and over 12 billion pages, spanning Scribd and Slideshare.
The corpus-wide backfill completed in a matter of months as Google Cloud scaled batch throughput to meet the timeline. Native PDF input let more than 99% of the corpus be processed as-is without building an OCR, rendering, or screenshotting pipeline.
Batch prediction's 50% discount versus interactive pricing made LLM classification economically viable at that scale.