Back Issues/Search Home → Calendar → Archive → RSS → Subscribe → Current Issue → Popular →

All issues › Volume 342, Issue 4 › IT Vendor News › Google

Scribd, Inc. Classifies More Than 400 Million Documents with Gemini Batch Inference on Gemini Enterprise

Google, Thursday, September 24th, 2026

Scribd used Gemini's PDF understanding and Gemini Enterprise batch inference to classify over 400 million user-uploaded documents.

Scribd used Gemini Enterprise's native PDF understanding and batch prediction to run trust-and-safety classification across its entire user-generated content corpus of more than 400 million documents and over 12 billion pages, spanning Scribd and Slideshare.

The corpus-wide backfill completed in a matter of months as Google Cloud scaled batch throughput to meet the timeline. Native PDF input let more than 99% of the corpus be processed as-is without building an OCR, rendering, or screenshotting pipeline.

Batch prediction's 50% discount versus interactive pricing made LLM classification economically viable at that scale.

more →  ·  More from Google →