How does the Siberson Veriket Data Discovery scan engine work?
Siberson Veriket Data Discovery uses a parallel, throttle-aware scan engine that can be tuned for any scale of environment agently or agentless. The engine ingests data from connectors and endpoint agents, runs every artifact through the content inspection pipeline, and writes risk-scored findings to the central inventory.
| Engine Capability | Description |
|---|---|
| Parallel scanning | Configurable worker count and concurrency per data source — scale from a single laptop to thousands of file shares without serializing |
| Incremental scans | After the initial baseline, subsequent runs re-process only changed content (modified-since timestamps, content-hash diff) |
| Throttling & QoS | CPU, memory, and I/O caps per scanner; off-hours scheduling; bandwidth limiting on remote shares |
| Resumable scans | Scans interrupted by network or power events resume from the last completed checkpoint |
| 400+ file types | OOXML, legacy Office, PDF, ODF, RTF, source code, archives (recursively), email containers, image OCR |
| OCR pipeline | Tesseract-based OCR for image-embedded text in scanned PDFs, screenshots, faxes — language packs include Turkish |
Last updated: 2026-04-29