Reading Images with Vision AI
Welcome to the Course 🚀
Welcome to AI Images for Workplace Communication. Images show up all over UX work: a usability-test session screenshot capturing exactly where a participant got stuck, a concept mockup or moodboard for a new flow, a hero slide behind the opening of a design review, a competitor's checkout screen you're pulling apart in a teardown, a screenshot headed into the research repository, a chart on an analytics dashboard. Generative AI now lets you describe, create, and edit those images in seconds — which means the bottleneck has shifted from "can I make this?" to "should I ship this, and does it say what I think it says about our users?" This course gives you the moves to handle both ends responsibly.
Course readiness: use only synthetic, stock, or explicitly approved non-sensitive images for every practice. Never upload a real participant's face or voice, personally identifying research data, unreleased designs, or confidential competitor material into a general AI tool. If you don't have image-capable access, you can still complete the work by writing and reviewing the prompts and briefs against the frameworks, using described or stock sample images instead of generating your own.
By the end of this lesson, you'll be able to:
- Request structured image descriptions from vision AI for a specific design purpose like research-repository alt text, research synthesis, usability-finding capture, or competitor-UI interpretation.
- Separate what an image actually shows from what AI (or you) is guessing about it, so an inference never becomes a fact in a research finding or a stakeholder update that sends engineers chasing the wrong fix.
- Turn a clean description into a usable design deliverable like accessible alt text, research-repository notes, accessible documentation, or a stakeholder update.
Before any AI-assisted design task, keep the Responsible AI Use Checklist in mind: approved tools only, safe inputs (never paste confidential product plans, roadmaps, or unreleased research), privacy & consent (remove or de-identify participant, teammate, and customer details before an image goes anywhere near a general AI tool), confidentiality, IP & copyright, and likeness rights. When in doubt, use sanitized, fictional, or permission-cleared artifacts — and remember an AI chat is not a system of record for research.
Asking Vision AI for the Right Kind of Description 🔍
Vision AI is the family of models that can "look at" an image you give it and describe it in words. The trap is treating it like a one-size tool. A generic prompt like "what's in this image?" gives you a generic answer: a wall of detail, half of which you don't need, with no structure you can paste into an alt-text field, a synthesis note, or a stakeholder update.
The fix is to name the purpose upfront. Different design jobs need different shapes of description:
- For alt text (the text a screen reader announces for an image in the research repository or team wiki), you want a tight, neutral sentence focused on visible content.
- For research synthesis, you want a plain inventory of what's literally on a usability-test screenshot before you draw any findings from it.
- For usability-finding capture on something like a screenshot of a participant stuck at checkout, you want the visible screen state separated cleanly from any guess about cause, attention, or intent.
- For chart interpretation on a screenshot of an analytics dashboard or funnel, you want axes, units, the trend, and the highest/lowest values called out.
A reliable prompt pattern looks like this: Describe this image for [purpose]. Use [format]. Focus on [what matters]. Skip [what doesn't].
For a usability-test screenshot headed to the research repository, you might write: Describe this image for alt text in our research repo. Use one neutral sentence plus a short bullet list. Focus on the UI elements and text actually visible on screen. Skip any guesses about why the participant paused, whether they noticed a control, or how they felt. You've now turned a vague request into something you can act on.

Telling Observation Apart from Interpretation ⚖️
Here is the single most important habit in this unit, and it's where research gets burned: vision AI will hand you observations and interpretations mixed together, in the same confident tone — and an interpretation that slips into a synthesis note can become a "finding" about cause that triggers an unwarranted redesign nobody can actually defend.
An observation is something visible in the pixels: "A red message reading 'Invalid entry' appears below the email field." An interpretation is a guess about meaning or cause: "The form validation blocked them from continuing." An inferred intent or assumption goes further: "The user didn't notice the primary button." Only the first kind is safe to build on. The other two are stories the model wrote from patterns in its training data, not evidence from this specific screenshot — and in a research note, "they didn't see the button" reads as a conclusion about attention and cause that a single frame can't prove.
Crucially, these three levels should lead to very different actions in UX work:
- Observations are verified facts you can act on directly: document them in research repository notes, alt text, or bug reports as objective visible evidence.
- Interpretations (guesses about cause or system behavior, like "the validation blocked them") are hypotheses to investigate: check session recordings, error logs, or telemetry data to confirm whether that was the actual technical cause before writing an engineering ticket or proposing a redesign.
- Inferred intent (guesses about user cognition or motivation, like "they didn't notice the button") cannot be determined from a static image at all: treat them as questions for follow-up user interviews or usability testing, and never base design decisions or stakeholder conclusions on inferred intent alone.
Watch for tell-tale words and phrases: "appears to," "looks like," "seems to," "suggests that," "clearly." Those are interpretation flags, as is any line that names a feeling, an intent, a cause, or a claim about "most users" the image can't actually prove.
Before we go further, the boundary that governs this whole skill. You never run a real participant's raw session — their face, their voice, their personal details — through a general AI tool. That kind of research material is only ever handled once it's de-identified or permission-cleared under your team's research policy, and interpretations are stripped out before anything reaches the design and engineering team as fact. For practice, you work from a de-identified training screenshot that stands in for a usability-test frame, so you can build the skill safely.
Here's how that filtering sounds when peer reviewer Dan catches it in conversation with Jessica, a product designer:
- Jessica: The vision AI says it's "the user didn't notice the 'Place Order' button, which is why they didn't finish." That's useful for the finding, right?
- Dan: What does the screenshot actually show?
- Jessica: A checkout form with a red "Invalid entry" message under the email field, and the Place Order button greyed out.
- Dan: Then say that. "Didn't notice the button" and "that's why they didn't finish" are a story — a claim about attention and cause. If we put that in the synthesis, we've decided why they stalled from one frame we can't back up.
- Jessica: Fair. I'll keep the note to the visible state and flag that cause isn't determinable from the screenshot.
Notice Dan isn't rejecting AI, he's rejecting AI-shaped fiction dressed up as fact. Your job is to be the human filter between the description and whatever the team synthesizes or acts on next — especially when a finding could send engineers rebuilding the wrong part of the flow.
Turning a Clean Description Into Something Useful 🛠️
Once you've stripped a description down to observations, you can shape it into the deliverable you actually need — but the deliverable depends entirely on which workflow you're in, and observation and interpretation must never be blurred.
For non-sensitive, approved images (a stock workspace photo, a de-identified UI screenshot cleared for the wiki), a clean observation set becomes alt text: short (aim for under 125 characters), neutral, and focused on visible content so teammates using screen readers get a faithful description. Alt text is concise and usually one sentence. Repository notes may include alt text plus visible-detail bullets. Synthesis notes can be more detailed, while separating observations from interpretations. It can also become accessible documentation or research-repository notes.
For research synthesis, the same observation discipline produces a stakeholder update that names what's visible, states explicitly what is not determinable from the screenshot (why the participant paused, whether they noticed a control, whether other participants hit the same thing, whether it reproduces), and recommends one concrete next step like "review the session recording at that timestamp before adding it to the findings backlog." The discipline is the same in every format: say what you can see, name the gaps out loud, and never let a confident sentence become a finding about your users just because it sounded authoritative.
The throughline of this unit: vision AI gives you a draft, not a verdict, and the value you add is sorting signal from story before anything reaches a stakeholder as fact. With that in mind, the next step is a live conversation: you'll walk a peer reviewer through an AI description of a de-identified usability-test screenshot and defend, line by line, which sentences are observations you can act on and which are interpretations you need to strip.
