Article

Research data classification guides licensing and access rights

What makes a piece of research data count as research data, and why does its type affect access and licensing? The answer matters because data that looks simple on the surface can carry different rights, different reuse rules, and different limits on sharing.

Research data is information gathered or made in the course of research to support or test a finding. It may be born digital, but it is not always digital. A lab notebook, a field diary, or a paper record can all count if they are part of the research record. That broad definition matters in libraries and archives because access is never only a technical issue. It is also a rights issue.

The first step is basic classification. In most research settings, data is grouped into two large types: qualitative and quantitative. Qualitative data describes qualities, meanings, or experiences. Quantitative data deals with numbers, counts, and measurements. The split is simple at a glance, but it shapes how the data is stored, searched, shared, and licensed.

Qualitative data often comes from interviews, focus groups, written notes, photographs, or observations. It may be text-heavy, image-based, or mixed. Because it often contains people’s words, faces, or stories, it can raise privacy and consent questions very fast. A dataset with interview transcripts may be useful to scholars, but access may be restricted because the material can identify a person or reveal sensitive details.

Quantitative data is usually numerical. It may come from surveys, tests, counts, measurements, or administrative records. It is often easier to sort and analyze at scale. Yet “numerical” does not mean “free of limits.” A table of figures can still be tied to contracts, institutions, or people. A dataset can be open in form and still closed in use.

This is where classification guides licensing and access rights. Once data is labeled and described, a curator, publisher, or repository can decide which rights apply. Some data can be shared openly. Some can be viewed only under a license. Some can be used only after approval. The classification does not create the rights by itself, but it helps show which legal and ethical questions need answers.

A plain example makes this easier to see. Imagine a researcher creates two files in one project. One file holds survey counts about reading habits in a town. The other holds interview notes from the same project, where people explain why they read or do not read. Both are research data. Both may support the same study. But the survey file may be easier to share in a public repository, while the interview notes may need more protection because they contain personal speech and context.

That difference affects more than access. It affects metadata, too. Metadata is the descriptive information that explains what a file is, who made it, when it was made, and how it may be used. Good metadata can state the data type, the format, the collection method, and the access terms. Without that information, a user may not know whether the item can be reused, cited, downloaded, or viewed at all.

Licensing is the formal part of that picture. A license tells users what they may do with the data. It can describe copying, reuse, redistribution, or adaptation. In library work, this is one of the places where plain description matters most. If the license is vague, the user does not have clear rights. If the license is clear, the user knows the boundary before they act.

Access rights are not the same as ownership. A researcher may hold a dataset and still restrict access for a time. A repository may host a file and still limit download or reuse. A user may be able to read a record but not republish it. These are different layers, and they should not be blurred together. Confusion here causes real problems, especially when people assume that open viewing means open reuse.

I often treat the label on the data as the first clue, not the final answer. “Qualitative” or “quantitative” tells me how the material is shaped, but not whether it is open, restricted, or licensed for reuse. To answer that, I look for the rights statement, the repository terms, and any notes on consent or confidentiality. If those are missing, the record is incomplete, even if the file itself is rich.

Search also changes by data type. Quantitative data may be easier to find through field names, date limits, or subject tags. Qualitative data often depends more on careful description, because a transcript or image set may not be searchable in the same way as a spreadsheet. That is not a flaw. It is a feature of the source. The search method has to match the shape of the material.

Classification also helps preserve context. A dataset stripped of its type, method, and license becomes risky to use. A user may mistake a working file for a published one. A repository may present a dataset without saying whether it is final, partial, or revised. In that state, access exists, but understanding does not. Research use depends on both.

This is why I value plain documentation so highly. A digital source is only useful when its coverage, search tools, and limits are stated clearly. For research data, that means the type of data, the terms of use, and the restrictions should all be visible in one place. If they are not, the source may still exist, but it is not fully usable in a scholarly sense.

The practical lesson is simple. Research data classification is not an abstract label. It is the hinge between description and permission. Once a dataset is identified as qualitative or quantitative, the next questions follow in order: what does it contain, who can see it, what may be done with it, and what limits are already built in.

After that, a reader can judge the record with more care. That is the skill this lesson builds. It lets the reader see how data type, licensing, and access rights fit together, and how a repository or database should explain them before trust is given. The Source List aims for that same plain view: one digital source worth knowing, one search tip, and one honest limitation.