Data classification is the process of categorizing data based on its sensitivity, value, and the level of protection it requires. Classification is a foundational element of both data governance and information security.
Why Classify Data
Not all data needs the same level of protection. A public press release and a confidential employee record should not be treated the same way. Classification creates a framework for applying appropriate controls to each type of data — and for communicating those requirements clearly to everyone who handles data.
Common Classification Levels
Public
Data that can be freely shared with anyone. Open data published on government portals is public. Marketing materials are public. No special handling required. Publishing public data openly supports transparency and civic engagement.
Internal
Data intended for use within the organization but not for public release. Internal reports, operational data, and general business information. Should not be shared externally without authorization. Most day-to-day business data falls into this category.
Confidential
Sensitive data that could cause harm if disclosed inappropriately. Customer personal information, financial data, and strategic plans. Requires access controls, encryption, and careful handling. Employees who handle confidential data should receive specific training.
Restricted
Highly sensitive data with strict access controls. Medical records, legal documents, and security-sensitive information. Access limited to specific authorized individuals. Breaches involving restricted data typically trigger mandatory notification requirements under privacy law.
Classification is a form of metadata: it describes a property of the data (its sensitivity level). Classification labels should be stored as metadata alongside the data and used to drive automated access controls and handling requirements. When classification is embedded in metadata, systems can enforce rules automatically rather than relying on individual judgment.
Classification in Practice
A practical classification exercise: list your organization's key datasets, assign each a classification level, and document the handling requirements for each level. This simple exercise often reveals that sensitive data is being handled as if it were internal, or that public data is being over-protected at unnecessary cost.
Key Takeaways
- Classification matches protection level to data sensitivity
- Four common levels: public, internal, confidential, restricted
- Classification is metadata — store it with the data
- Automated enforcement is more reliable than manual judgment
Learn how to build an inventory of your data assets in Cataloguing.