What does redaction mean in the context of PDFs?
Redaction sounds simple: place a black rectangle over sensitive text or images, done. But in practice, this is a frequent and serious mistake. A black rectangle placed over text visually conceals it, but the original text remains in the document and can be extracted using a text editor or simple copy-paste.
True redaction means: the original content is permanently removed from the document. After this, there is no way to recover the redacted content, not even through technical means.
Why redaction matters, the GDPR perspective
Data protection regulations require organizations and public bodies to protect personal data. When documents are shared, with third parties, business partners, in legal proceedings, all information not intended for the recipient must be properly redacted.
Common use cases:
- Contracts: redact names, addresses, bank details of other parties
- Personnel files: hide health data or salary information from certain recipients
- Official documents: case numbers, personal identifiers, addresses in file disclosures
- Invoices: bank details, personal addresses when forwarding
- Medical records: diagnoses, treatment details
The most common mistake: covering instead of redacting
Time and again, organizations share documents with a black bar placed over content, while forgetting that the text underneath is still in the document. This mistake has led to serious data breaches in the past, because recipients could extract the text despite the apparent redaction.
An example: a PDF with a black text box over an address. In the PDF viewer, you only see the black rectangle. But if you copy all page content and paste it into a text editor, the address appears anyway.
How proper redaction works
With redact PDF, selected areas are permanently removed from the document. The process works in two steps:
- Mark areas: text or image areas to be redacted are selected in the document
- Apply redaction: when saving, the marked content is permanently removed from the document and replaced with a black rectangle
After this process, there are no hidden text layers remaining. The redacted content is irreversibly removed.
Text search and bulk redaction
If a specific term, for example a name or account number, appears in multiple places in a document, manually marking each instance would be tedious. Professional redaction tools therefore allow searching for specific terms and applying redaction to all occurrences.
With automated redaction, a manual review is always recommended afterwards: was every relevant instance captured? Are there additional instances not found through search (for example, names in graphics)?
What else to consider after redaction
Beyond the actual content, PDFs can contain additional sensitive information that is often overlooked:
- Metadata: creator, edit date, software, viewable in document properties
- Comments and annotations: may contain information not intended for the recipient
- Embedded files: some PDFs contain embedded attachments
Anyone wanting to additionally secure a document after redaction can use protect PDF to add a password, disable editing and allow only reading.
Black boxes over image areas
For photos or scans, redaction works slightly differently: since the content is already an image, it is technically sufficient to cover the image area with an opaque black rectangle, as long as the image is a raster image, not a layered element. When in doubt, use a proper redaction process to be safe.
Conclusion
Proper redaction in PDFs is not a luxury but a data protection necessity. Anyone sharing sensitive documents must ensure that redacted content is truly and permanently removed, not just visually concealed. The right tool makes the difference between a secure document and a potential data protection risk.