Welcome to the Course

As a Data Governance Analyst, you already feel the pull of two forces every week: colleagues who want to use data fast, and a quiet worry about whether a given use is actually safe. This course gives you a calm, repeatable way to handle that tension without needing to be a lawyer, a security engineer, or a privacy officer. The whole idea is risk triage for non-specialists: a steady habit you apply to any data request so you can decide what you safely can, document it, and escalate the rest to the right expert.

By the end of this course, you'll be able to:

  • Triage any data risk with a repeatable pattern: identify the risk, gather context, consult the right partner, document a recommendation, and escalate when needed.
  • Classify data as public, internal, confidential, or restricted using plain business-risk criteria.
  • Apply core privacy principles to narrow over-broad data requests without blocking the work.
  • Coordinate with security partners using the CIA triad, supplying the inputs they need while leaving implementation to them.
  • Support retention and disposal decisions that balance business, legal, operational, and historical needs.
  • Spot ethical and bias risks, such as proxy variables and unfair exclusion, and recommend safeguards.

This first unit builds the foundation everything else rests on: how to triage a risk, how data sensitivity shifts across its lifecycle, and how to classify a dataset by how much harm its misuse could cause.

Triaging Risk Across the Data Lifecycle

Before you classify anything, you need a mental routine you can run on autopilot. That routine is the risk-triage pattern, and it has five moves. First, identify the risk: ask "what could go wrong if this data were seen, changed, or shared by the wrong person?" Second, gather context: who wants it, why, for how long, and where will it go? Third, consult the right partner when the risk is beyond your seat. Fourth, document a clear recommendation so the decision is visible and reusable. Fifth, escalate when the stakes are high, meaning regulated data, external sharing, or anything you are genuinely unsure about.

Notice your job here. You are not the final approver on the hardest calls. You are the person who triages quickly, handles the routine cases, and routes the serious ones to legal, privacy, security, or compliance with the facts already organized. That framing keeps you fast without making you reckless.

The second piece of this foundation is that risk is not fixed. The same dataset carries different expectations depending on where it sits in its data lifecycle: creation, storage, active use, archive, retention, and disposal. A fresh customer record in active use needs to be accurate and available to the teams working it. That same record, years later in archive, should be locked down tighter, shared with far fewer people, and eventually disposed of responsibly. So when you triage, always ask which stage the data is in, because the stage changes what "responsible handling" looks like. Data you would happily share during active use might be something you minimize or delete entirely once its purpose is done.

Classifying Data by Sensitivity

With the triage habit in place, you need a shared language for "how sensitive is this?" That language is the Four-Tier Classification. Public data causes no harm if disclosed, like a published price list. Internal data causes low harm and should stay among employees, like a routine team roster. Confidential data would cause significant harm if it leaked and belongs on a need-to-know basis, like a customer list joined with purchase history. Restricted data would cause severe harm, often because it is regulated or highly sensitive, like payment card numbers or health information.

The four-tier data classification as an ascending scale of harm: Public causes no harm if disclosed, Internal causes low harm and stays among employees, Confidential causes significant harm and is need-to-know, and Restricted causes severe harm and is often regulated

The trick most newcomers miss is that sensitivity is not just about the fields themselves, it is about audience and combination. A name alone might be internal. A name combined with what someone bought, and then sent to an outside vendor, tells a revealing story and travels outside your walls. That combination plus that audience pushes the tier up.

  • Dan: It's just our customer list, can I send it over to the agency?
  • Victoria: Names alone might be internal, but you're attaching purchase history too, right?
  • Dan: Yeah, so they can tailor the campaign.
  • Victoria: That combination tells a story about each person, and it's leaving the company. That moves it to confidential, need-to-know.
  • Dan: So I can't share it?
  • Victoria: You can, just on a need-to-know basis with handling conditions, not a broad blast.

Watch how the answer is never a flat "no." It is "yes, and here's how." That stance is what keeps colleagues bringing requests to you instead of routing around you.

Applying the Decision Tree and Setting Handling Expectations

To make your classification consistent rather than gut-feel, walk the Classification Decision Tree every time. Identify the data type, determine its sensitivity and any obligations, assess the potential harm if it were misused, assign a tier, set handling expectations, and review when the context changes. The harm-if-misused step is the heart of it: you are reasoning about consequences, not just labeling fields.

Once you land on a tier, that tier drives concrete handling expectations across four areas: who may access it, how it can be shared, how long it is retained, and how it is disposed of. The higher the tier, the tighter each of those becomes. A confidential dataset shared externally, for instance, should go to named recipients only, under an agreement, with a defined end date and confirmed deletion afterward, not a "keep it forever, share as you like" arrangement.

Here is the boundary to hold firmly: you recommend, you do not finalize. When a request touches regulated data, sensitive personal information, external sharing, or high-impact disposal, your move is to document your recommendation and escalate to a legal, privacy, security, or compliance specialist who makes the final call. You bring them a clean, triaged picture instead of a vague worry, which makes their job faster and your judgment trusted.

The single takeaway of this unit: sensitivity is about the harm misuse could cause, given the audience and the lifecycle stage, and your role is to triage and recommend a tier, not to rule on the hardest cases alone. Next, you'll run a quick set of real-world scenarios and decide, on the spot, whether each one is public, internal, confidential, or restricted. Treat it as pattern practice: for each one, ask "who could see this, and how badly would it hurt if the wrong person did?" and let that answer point you to the tier.

Sign up
Join the 1M+ learners on CodeSignal
Be a part of our community of 1M+ users who develop and demonstrate their skills on CodeSignal