Siberson
Partnership Contact Request a Demo

How does the Siberson Veriket Data Discovery scan engine work?

Siberson Veriket Data Discovery uses a parallel, throttle-aware scan engine that can be tuned for any scale of environment agently or agentless. The engine ingests data from connectors and endpoint agents, runs every artifact through the content inspection pipeline, and writes risk-scored findings to the central inventory.

Engine Capability Description
Parallel scanning Configurable worker count and concurrency per data source — scale from a single laptop to thousands of file shares without serializing
Incremental scans After the initial baseline, subsequent runs re-process only changed content (modified-since timestamps, content-hash diff)
Throttling & QoS CPU, memory, and I/O caps per scanner; off-hours scheduling; bandwidth limiting on remote shares
Resumable scans Scans interrupted by network or power events resume from the last completed checkpoint
400+ file types OOXML, legacy Office, PDF, ODF, RTF, source code, archives (recursively), email containers, image OCR
OCR pipeline Tesseract-based OCR for image-embedded text in scanned PDFs, screenshots, faxes — language packs include Turkish

Last updated: 2026-04-29