Data classification is the process of categorizing data based on its sensitivity, value, and the level of protection it requires. Classification is a foundational element of both data governance and information security.

Why Classify Data

Not all data needs the same level of protection. A public press release and a confidential employee record should not be treated the same way. Classification creates a framework for applying appropriate controls to each type of data — and for communicating those requirements clearly to everyone who handles data.

Common Classification Levels

Public

Data that can be freely shared with anyone. Open data published on government portals is public. Marketing materials are public. No special handling required. Publishing public data openly supports transparency and civic engagement.

Internal

Data intended for use within the organization but not for public release. Internal reports, operational data, and general business information. Should not be shared externally without authorization. Most day-to-day business data falls into this category.

Confidential

Sensitive data that could cause harm if disclosed inappropriately. Customer personal information, financial data, and strategic plans. Requires access controls, encryption, and careful handling. Employees who handle confidential data should receive specific training.

Restricted

Highly sensitive data with strict access controls. Medical records, legal documents, and security-sensitive information. Access limited to specific authorized individuals. Breaches involving restricted data typically trigger mandatory notification requirements under privacy law.

Classification as Metadata

Classification is a form of metadata: it describes a property of the data (its sensitivity level). Classification labels should be stored as metadata alongside the data and used to drive automated access controls and handling requirements. When classification is embedded in metadata, systems can enforce rules automatically rather than relying on individual judgment.

Classification in Practice

A practical classification exercise: list your organization's key datasets, assign each a classification level, and document the handling requirements for each level. This simple exercise often reveals that sensitive data is being handled as if it were internal, or that public data is being over-protected at unnecessary cost.

Key Takeaways

  • Classification matches protection level to data sensitivity
  • Four common levels: public, internal, confidential, restricted
  • Classification is metadata — store it with the data
  • Automated enforcement is more reliable than manual judgment
Next Step

Learn how to build an inventory of your data assets in Cataloguing.

← Data Dictionaries Cataloguing →