Choosing Where AI Helps at Work
Choosing Where AI Helps at Work ðŸ§
Now that you can tell the AI capabilities apart, the next move is harder and more useful: deciding which of your actual design tasks belong in an AI workflow and which ones absolutely don't. Choosing the right starting point matters. A use case that looks exciting but is hard to verify—like AI-written accessibility annotations—can drain time and create compliance risk, while overlooking practical everyday tasks can leave real value untapped. This unit gives you a simple lens to sort your week and choose a first win you can defend.
By the end, you'll be able to:
- Sort recurring UX design work into communication, research, planning, and analysis tasks.
- Evaluate AI candidates using Value, Risk, Repeatability, and Human Review.
- Select a low-risk pilot with a clear success criterion.
Mapping Your Work Into AI-Assistable Categories 🪣
Start by looking at a normal week and asking: what kind of design work keeps showing up? Most AI-assistable tasks fall into three practical buckets.
| Bucket | What It Means | Common UX Examples |
|---|---|---|
| Communication | Writing or shaping information for an audience | Weekly design-status updates, stakeholder notes, release-note UX copy, design-handoff summaries |
| Research & Analysis | Gathering, digesting, and making sense of information | Summarizing usability-test notes, scanning competitor UX patterns, tagging themes in research notes, comparing design options against heuristics |
| Planning | Structuring work before it happens | Drafting research-recruiting screeners, building design-review agendas, creating design-handoff checklists |
Why bother sorting? Because each bucket has a different risk profile and a different verification cost. A draft design-status update in Communication may need only a quick accuracy check against your project board. An insight or heuristic evaluation in Research & Analysis may need a specialist to verify it. If you don't separate these task types, you may treat every AI output the same way, which is exactly how a confident-but-wrong WCAG guidance line ships into a product and creates an inaccessible—potentially non-compliant—experience.
Try This: Take ten minutes this week and list five recurring design tasks under each bucket. That list becomes the candidate pool for everything that follows.
The AI Task Fit Screen 🤖
Once you have candidates, run each one through the AI Task Fit Screen, four quick checks that tell you whether a task is a fit, a maybe, or a hard no.
Value asks whether the payoff is worth it. Look for tasks where AI can meaningfully reduce time, improve consistency, or help you get unstuck. If AI shaves three minutes off something you do twice a year, the math doesn't work.
Risk asks what a bad output could cost—and what a careless input could expose. A clumsy phrase in an internal design-status update is recoverable; a wrong WCAG success-criterion reference in an accessibility annotation, a misattributed research finding, or a fabricated competitor feature claim is not. Risk also runs in the other direction: never paste unreleased product roadmaps, confidential design specs, proprietary user-research data, or session recordings into an AI tool unless it's been explicitly approved for that use, since anything you submit may be stored or used to train the model. Run the Responsible AI Use Checklist before you start, applying the same checkmarks introduced in Unit 1:
- Approved Tools: Use only tools approved for your design work.
- Safe Inputs & Confidentiality: Keep confidential product roadmaps, proprietary user-research data, and unreleased design specs out of prompts.
- Privacy & Consent: Ensure participant, employee, and customer details are removed or de-identified; confirm consent before using research assets.
- IP, Copyright & Likeness Rights: Never feed in or reproduce protected third-party assets, brand logos, or recognizable likenesses without permission.
- Human Review & Verification: Verify all factual claims before they leave the team, especially when evaluating risk in the task fit screen.
- System of Record: Treat AI sessions as drafting environments, not systems of record.
Repeatability asks whether the task happens often enough to justify the setup. Recurring work—like a weekly design-status update or a batch of research-screener first drafts—is usually a stronger candidate because you can reuse the prompt, refine the workflow, and compare results over time.
Human Review asks whether a person can actually sanity-check the output in reasonable time. If you can't tell whether a WCAG criterion was cited correctly or whether a research summary accurately represents what participants said, you can't use the tool safely, no matter how polished the writing sounds. And for anything that touches accessibility compliance, participant privacy, or legal obligations, "human review" may mean escalation to an a11y specialist or your legal team—not just you—which changes how fast and how cheaply a task can really be reviewed.
The trap most designers fall into is anchoring on Value and ignoring the other three. High value plus high risk plus weak human review is not a pilot, it's a future compliance incident.
Here's what that sounds like in practice.
- Natalie: I want our first AI pilot to be writing the accessibility annotations straight from our design specs. It would save the design and eng team a huge amount of time.
- Jake: Value's real, agreed. But let's run it through the screen together. What's the cost if it gets a WCAG success criterion wrong?
- Natalie: Could be an inaccessible product, maybe even non-compliant. That's not a good look.
- Jake: And can someone without an accessibility specialist review the output and catch that?
- Natalie: Honestly, no. We'd still need an a11y expert to read every line before it went into the spec.
- Jake: So Value's high, Risk's high, Human Review's weak unless we loop in a specialist every time. Worth exploring with an a11y partner someday, but not our first pilot.
Notice Jake didn't kill the idea—he named the framework, walked the checks, and let the score speak. That's the move when someone anchors on a flashy candidate.
Picking Your First Pilot 🚦
Your first AI workflow should be deliberately boring: a task that scores well on all four checks, with low risk if it goes sideways. Recurring internal communication is a classic fit (the weekly design-status update, a research-screener first draft, a design-handoff summary) because the value is real, the risk is contained, the task repeats regularly, and you can eyeball the output in under a minute.
Then commit to a success criterion before you start, not after. Something concrete you can measure, like "cuts my design-status drafting time in half across three consecutive weeks with no factual corrections needed." Vague criteria like "saves time" let you talk yourself into a pilot that didn't actually work.
The takeaway for this unit: pick the task, not the tool, and let the AI Task Fit Screen tell you which task earns the pilot slot. The next step is a live conversation where you'll pressure-test five candidate design tasks against the Screen with a Head of Design who's anchored on the flashiest option. Bring the framework by name and walk the checks out loud—that's where the skill actually gets tested.
