Supported File Formats for Scanning
Forcepoint DSPM extracts text and metadata from files to support data classification and DLP policy evaluation. This topic describes which file formats are supported for scanning.
Overview
Forcepoint DSPM extracts text and metadata from files to support data classification and DLP policy evaluation. This topic describes which file formats are supported for scanning, the file size limits that apply, how archives and emails are handled, and how the platform prioritises files when multiple items are queued for scanning.
File type detection
The platform determines file type by analysing file content, not by reading the file extension. This means that renaming a file does not affect how it is detected or processed — the actual content is always used to identify the true format and apply the correct extraction method.
What is extracted
For each supported file, the platform extracts the following independently:
- Body text — raw text extracted from the file content, normalised to Unicode, and passed to the classifier.
- Document metadata — structured fields such as author, title, and creation date, extracted separately.
- Images — standalone image files are processed via OCR where applicable (see OCR section below).
All supported file formats support full content extraction.
Supported file formats
The table below lists all supported file formats by category. DSPM may be able to scan file types not listed here.
| Format category | Supported file extensions |
|---|---|
| Word Processing | ABT, ABW, ACT, ASD, AX, BOXNOTE, CHI, CN, DC, DOC, DOCM, DOCX, DOTM, DOTX, FNX, FR, GL, HLP, HW, HWP, HWPX, IC, ICHAT, INF, IWA, JTD, JW, KWD, LEX, LSM, LYX, MAIL, MM, MSW, MWK, NB, NN, ODT, ONE, ONETOC2, OTT, PFS, PPP, PTX, PUB, RFT, RTF, SAM, SDW, SGL, SMT, SPW, SS, TSCPROJ, TW, TXT, UNI, UOF, UOT, VOR, VW4, W40, WBK, WDM, WOP, WP, WP4, WPD, WPF, WPL, WPS, WRI, WS, WSD |
| Spreadsheets | 123, AS, CSV, DIF, ET, KSP, ODS, OTS, PSV, QPW, S30, S40, SDC, SS, STC, SXC, TAB, TDS, TDSX, TMS, TPS, TSV, TSX, TWB, TWBX, UOS, WB1, WB2, WB3, WK1, WK3, WK4, WKS, XLAM, XLR, XLS, XLSB, XLSM, XLSX, XLTM, XLTX |
| Presentations | CDPZ, CHRT, DPS, KARBON, KEY, KPR, MMP, MOM, NUMBERS, ODP, OTP, PAGES, PKPASS, PKPASSES, PMOM, POTM, POTX, PPAM, PPSM, PPSX, PPT, PPTM, PPTX, PRC, PRE, PRZ, SDA, SDD, SHW, SXD, SXI, UOP |
| PDF and eBooks | AZW, AZW1, EPUB, FB2, FDF, ION, KFX, OXPS, PDF, PFILE, PJPG, PPDF, PTXT, TPZ, TR, XFDF, XPS |
| Email and Calendar | EML, EMX, ICS, MAIL, MBX, MHT, MHTML, MSG, OFT, OLM, SMTP, VCF, VCS |
| Archives and Containers | 7Z, A, AAB, AD, AD1, AES, APK, AR, ARJ, BAK, BIN, BZ2, CAB, CER, CPIO, CRT, DAA, DEB, DER, DF1, DMG, DXL, EAR, EGG, EXP, GZ, HDR, IMAGE, IPA, ISO, JAR, KDB, KDBX, KS, LBR, LBS, LHA, LIB, LZH, MAFF, MAR, MFL, MSI, MSIX, MTF, NUPKG, O, ONEPKG, P7S, PACK, PAX, PEM, PUP, R00, R01, RAR, REV, RPM, SAR, SWD, SWF, SZ, TAR, TORRENT, UUE, VHD, WAR, WIM, XAPK, XPI, XZ, ZIP, ZIPX, ZSTD, ZXP |
| Images (OCR, English only) | $RH, AFF, APNG, BGA, BMP, COD, CT, DCM, DEEP, DIB, EMF, FLIF, GD, GD2, HEIC, HEIF, HFA, ICO, IM, IMG, J2K, JFI, JFIF, JP2, JPEG, JPF, JPG, JPM, JPWL, JPX, MIFF, NOFF, PAL, PGX, PKPASS, PKPASSES, PNG, PSP, PSPIMAGE, QPG, RAS, RH, RHM, RIX, RLE, RS, SCT, SCX, SUN, SVG, SYM, TDD, TIF, TIFF, USDZ, VIC, VICAR, WEBP, WMF, ZBR |
| Plain Text and Markup | AB, AJP, AMF, AML, ATOM, AXL, CDXML, CML, CSPROJ, ES3, FO, GHX, GLTF, GRA, GRAFFLE, GRF, GRP, GRXML, HPG, HPGL, HTM, HTML, IDF, IFCXML, JNLP, JSON, JVX, KML, KMZ, LNK, MARC, MD, METALINK, MIF, MO, MODS, MSS, MXL, MXML, NAB, ODF, ODG, OPF, OPML, OTG, OTH, OXT, PCL, PGML, PGN, PLS, PPD, PRN, PXL, RAML, RDF, RDL, RDOC, RDS, RSD, RSS, SAMI, SBML, SCC, SGML, SMIL, SRGS, SRL, SRU, SRX, SSML, TEI, TMX, TTL, TTML, VCXPROJ, VI, VTU, VXML, WML, WSF, XBAP, XBRL, XDF, XDP, XHT, XHTML, XML, XQM, XSL, XSLFO, XSLT, XSPF, XUL, YML |
| Source Code | ABAP, AGDA, AJ, ALS, AMPL, APL, APPLESCRIPT, ASC, ASN, AWK, B, BAT, BF, BMX, BRD, BRS, BSV, C, CAKE, CBL, CCP, CEYLON, CHPL, CL, CL2, CLJ, CLP, CLS, CLW, CMAKE, COB, COFFEE, COM, CP, CPP, CPY, CR, CREOLE, CS, CSD, CSS, CU, D, DART, DBL, DCL, DOT, E, EC, ECL, EL, ELM, EM, ERL, ES, F, FAN, FOR, FORTH, FS, FTL, FZ, FZB, FZBZ, FZP, FZPZ, G, GLSL, GML, GMS, GNU, GO, GOLO, GP, GRADLE, GRAPHQL, GRT, GS, GVY, H, HAML, HBS, HLSL, HPP, HS, HY, I7X, ICL, IDR, IJS, IK, INO, IPF, J, JAVA, JL, JQ, JS, JSX, KT, LAS, LASSO, LFE, LOL, LS, LUA, M, MAKE, ML, MONKEY, MOO, MS, MXT, NIX, NL, NLOGO, NSI, NU, NUT, OX, OXYGENE, OZ, P, PAN, PASCAL, PASM, PB, PDE, PHP, PICKLE, PIKE, PKL, PL, PLSQL, PONY, PP, PRO, PROLOG, PS1, PUP, PWN, PY, R, RB, REB, REBOL, RED, REXX, RING, ROBOT, RPY, RS, RSC, SAS, SC, SCAD, SCH, SCI, SH, SLS, SMI, SPC, SRGS, SSML, ST, STAN, STYL, SV, SWIFT, T, TM, TS, TTL, TXL, UR, URS, V, VB, VHOST, VIM, W, WAT, WEBIDL, X10, XQM, XTEND, YANG, YIN, YML, ZEP |
| CAD and Vector Graphics | 3DXML, AC, ASM, CATPART, CATPRODUCT, CDDX, CDR, CGM, CMX, CNC, CO, DAE, DGN, DRW, DWF, DWFX, DWG, DWS, DWT, DXF, EDF, FBX, FCSTD, FM, FRM, GBR, GDS, GDS2, GHX, GLP, GLTF, GXF, ICD, IFCXML, IGS, LWS, MA, MODEL, MPP, MPX, MSH, MTL, OBJ, PLY, PRT, RFA, RIB, RTE, RVT, SAT, SMP, STEP, STL, STP, STPNC, USDZ, VDX, VSD, VSDM, VSDX, VSSM, VSSX, VSTM, VSTX, WPG, X, X3D, XE, X_T |
| Audio and Video | 3G2, 3GP, AIF, AIFC, AIFF, APC, ASF, ASX, AU, AUX, CSP, DTS, DVF, F4P, F4V, FLW, GXF, HNM, ISS, M2T, M2TS, M3U, M4A, M4B, M4P, M4V, MMM, MOD, MOV, MOVIE, MP3, MP4, MPC, MPEGA, MPG, MQV, MSV, MTS, MV, MXL, NIST, OGM, PAC, QT, RSO, SAN, SND, SOL, SPC, STR, THMX, WAX, WMA, WMV, XA, ZIK |
| Database | ACCDB, ACCDT, AVRO, DBF, GDB, HL7, MDA, MDB, ORC, PARQUET, RDL, RDS, RRD, SAS7BDAT, VCX |
| GIS | AXL, BPW, CSF, GFW, ISO19139, J2W, JGW, KPLATO, LAS, LAZ, MAP, MPS, PGW, PRJ, TFW, WLD |
| Scientific and Data | CDXML, CML, DAT, FCS, FIG, GB, H4, H5, HDF, LBL, MAT, MATHML, MML, SAV, SBML |
| Subtitles and Captions | ASS, SAMI, SCC, SRT, SSA, TTML |
| Fonts | AFM, BDF, CPX, PFA, PFB, SFV |
File size limits
- Individual files: The maximum size for a single file submitted for scanning is 15 MB. This applies to the raw file size regardless of file type. Files that exceed this limit are skipped.
- Archives: The platform limits the total size of content extracted from an archive to 50 MB.
Archives (ZIP, RAR, TAR, and similar)
When an archive is scanned, the platform unpacks it and applies the 15 MB limit to each contained file individually. Files inside an archive that exceed 15 MB are skipped, but all other files continue to be processed. Nested archives are unpacked recursively under the same rules.
Emails with attachments (EML, MSG)
Email files are treated as containers. The email body is extracted first, then each attachment is processed individually. The 15 MB limit applies per attachment. If an attachment exceeds 15 MB it is skipped; the email body and remaining attachments are still processed.
How archives are classified
An archive is classified based on the aggregate of all files extracted from it. A policy match on any file inside the archive will be reflected in the archive's classification result.
Scanning priority
When multiple files are queued, the platform assigns priority based on the following rules. All queued files are eventually processed; priority only affects the order in which results become available.
- Files from DDR are prioritised over files from unstructured scanning.
- Among files of the same source type, smaller files are processed before larger files.
- Archives and images are assigned lower priority than other file types.
| File category | Priority |
|---|---|
| Files from DDR (text files ≤ 1 MB) | 1 (Highest) |
| Files from DDR (text files 1–5 MB) | 2 |
| Files from DDR (text files 5–10 MB); or files from DDR (images and archives ≤ 5 MB) | 3 |
| Files from DDR (text files > 10 MB); or files from DDR (images and archives > 5 MB) | 4 |
| Files from unstructured and structured scanning (excluding images and archives) — small (≤ 1 MB) | 5 |
| Files from unstructured and structured scanning (excluding images and archives) — medium (1–5 MB) | 6 |
| Files from unstructured and structured scanning (excluding images and archives) — large (5–10 MB); or images and archives ≤ 5 MB | 7 |
| Files from unstructured and structured scanning (excluding images and archives) — very large (> 10 MB); or images and archives > 5 MB | 8 (Lowest) |
OCR — extracting text from images
The platform supports Optical Character Recognition (OCR) for standalone image files such as PNG, JPEG, TIFF, BMP, and WEBP.
OCR is supported for English only. Documents containing image-based text in other languages will not have that image text extracted or classified.
When a file cannot be extracted
| Reason | Outcome |
|---|---|
| Unsupported format | File type is recorded and text extraction is skipped. The file status is set to EXTRACTION_IMPOSSIBLE. |
| Corrupt or truncated file | Partial extraction is attempted. If extraction fails entirely, the file status is set to EXTRACTION_IMPOSSIBLE. |
| Encrypted or password-protected | No content is returned. The file is identified as encrypted during the detection pass. The file status is set to EXTRACTION_IMPOSSIBLE. |
| File exceeds size limit | The file is skipped. The file status is set to FILTERED_OUT. |
| Extraction timeout | The operation times out. The file status is set to EXTRACTION_IMPOSSIBLE. |