Work Play About Contact

Product Design · UX Research · AI Platform

Microsoft Foundry
Models

Microsoft Foundry’s catalog had grown past 11,000 models, with no reliable way to find or compare them. Our team of five redesigned discovery around the decisions developers actually make: narrowing, evaluating, and committing. The redesign shipped into the live product.

My Role
User Experience Designer
Team
5 UX designers
Timeline
April to June 2026
Co-op
Microsoft CoreAI
Tools
Figma, UXtweak, Netlify, Claude, Cursor
ai.azure.com › Discover › Models
Explore the live platform
At a glance
01
The challenge
A catalog of 11,000+ AI models with no clear place to start, and no sign of which ones a developer could even deploy in their region.
02
What I did
On a team of five, I ran the usability research and led concept and IA, owning the two threads that shipped: region visibility in the filter rail and side-by-side compare. Both are live today.
03
The outcome
Both ideas shipped into the live Microsoft Foundry product (region in the filter bar, a side-by-side compare tray), and a small round-two test pointed to roughly 58% faster search-to-select.

The Problem

11,000 models,
and nowhere to start.

Developers hit 11,000 models with no way to narrow the list, and no signal of what would even run in their region.

Picture a developer who just wants to ship a feature this week. They open Microsoft Foundry, land on the Models page, and hit more than 11,000 models, each with different capabilities, pricing, context windows, and regional availability. Users kept telling us the same two things: help me narrow this list, and help me figure out where to start. Some picked a model only to hit a wall at deployment because it wasn't available in their region.

Our team of five spent three months redesigning model discovery end to end. I came to the project from the research side, after running a UX research study on Foundry.

The old Foundry Models page: a dense four-column grid where each model shows only a name and a task, and the filter rail has no way to narrow by region

Toggle between the old catalog and the shipped redesign. Before: a dense multi-column grid where every model shows little more than a name and a task, with no way to narrow by region, so a developer could commit to a model and only hit a wall at deployment. After: Region is promoted into the filter rail as a first-class way to narrow the catalog, in a cleaner, more scannable layout, with Compare a click away.

User Research

Developer interviews shaped
our direction.

Every method pointed to the same gaps: unclear terms, dense structure, and no way to compare.

We ran our own study: a heuristic evaluation of the live platform, moderated developer interviews, and a card sort.

🔍
Heuristic evaluation
  • Unclear terminology
  • Dense information structure
  • High cognitive load
  • Lack of onboarding support
🎙️
User interviews
  • Comparing models meant too much clicking back and forth
  • Lost their place navigating a large catalog
  • Users needed key model details earlier in the flow
🃏
Card sorting
  • Explored how users group and organize information
  • Revealed confusion around similar model tasks
  • Guided the creation of our information architecture

Who we designed for

Research pointed to three personas: Maya, a student developer, and Priya, a product manager weighing cost against capability. After aligning with Microsoft, we designed primarily for Alex, a startup developer who needs to move fast without picking a model that causes problems later.

Primary persona
👨🏻‍💻
Alex
Startup Software Developer

“I want to move fast without picking a model that creates problems down the road.”

Goal: find and select the right model quickly for a feature or prototype.
Pain: too many models shown at once; hard to find region-relevant options fast.

Sketches & Wireframes

Sketching our way
out of the problem.

Two decisions drove the redesign: where the region problem gets solved, and what a developer sees first.

Before touching Figma, we sketched. Paper let us argue about structure instead of pixels, which is where those two decisions got settled.

Hand-drawn paper prototype: a region blocker handled in onboarding, and a models selection flow with an AI helper and clearer descriptions

Paper prototype. Two ideas we carried forward: handling region up front so a developer never commits to a model they can't deploy, and a model-selection flow with an “ask AI to find the best model” helper and clearer descriptions to reduce overload.

Three ideas carried from paper into the mid-fidelity wireframes:

📍 Region & pricing visible right in the results 🔎 A more prominent, AI-assisted search 🧩 Clearer descriptions to reduce overload

Designing for the states that aren't the happy path

A catalog of 11,000 models is mostly edge cases. Three shaped the design as much as the main flow, each a decision with a real trade-off.

🌍
A model you can't deploy
Region-locked models stay visible with a clear availability flag, instead of vanishing and leaving users wondering why.
🔍
Filters that return nothing
A zero-result state names the filter that emptied the list and offers a one-tap way to loosen it, so a blank screen never reads as broken.
⚖️
Compare before it's full
The compare tray opens with a prompt and a running count, activating side-by-side only at two or more, so it never feels broken on first contact.

Three directions we killed

With a fixed IA and a tight scope, the hard part was deciding where not to spend the effort. The shipped answer looks obvious in hindsight; it wasn't. I explored three approaches that tested worse, and killing them is what made the final design defensible.

01
Let the system pick the model for you
A recommender that took your task and returned one “best” model, skipping the catalog entirely.
Why it died: developers wouldn't put a black-box pick into production. They wanted to see the trade-offs and make the call, so I moved the energy into compare, not auto-select.
02
Region as a badge on every card
Stamp each of 11,000+ model cards with its regional availability so nothing surprised you at deploy.
Why it died: at that scale it was visual noise and still didn't let you narrow. Region only pays off as a filter, so it moved into the rail.
03
A full spec sheet on each result
Put context window, pricing, latency, and modality on every tile so the list was fully informative.
Why it died: it made scanning 11,000 rows worse, not better. Detail belongs where the decision happens, so it moved into compare.

Mid-fidelity wireframes

With the structure settled, I moved into Figma to resolve layout and hierarchy. These mid-fidelity screens worked out where each decision lives; the high-fidelity, clickable prototype comes next.

Discover: a front door with intent

Instead of an 11,000-row wall, Discover opens with what you're trying to do: find, compare, try, or fine-tune a model.

Foundry Discover page: four task cards (find, compare, try, fine-tune), featured models, providers, and model collections

Models: filter, region, and compare in one place

The core of the redesign: a filter rail narrows 11,000+ models, region and pricing sit right in the results, and a compare tray collects models for a side-by-side decision.

Foundry Models page: filter rail on the left, results table with region, pricing, and latency, and a compare tray holding three selected models

Home: pick up where you left off

Returning developers land on recent models and quick tasks, so a half-finished decision doesn't mean starting the search over.

Foundry home screen: a welcome header, primary actions, a pick up where you left off row, and quick task shortcuts

High-Fidelity Prototype

The clickable, high-fidelity build.

Before it shipped, the redesign lived here: a full high-fidelity prototype, embedded and clickable. Explore Discover, the Models page, filtering, and side-by-side compare, then see the shipped solution next.

Interactive Walkthrough

ms-foundry-prototype.netlify.app

Solution

Two ideas, now shipped.

Two ideas I drove shipped into Microsoft Foundry, and both are live in the product today. The first promoted Region into the filter rail, so developers narrow the catalog by where a model can actually deploy, before they commit rather than after they hit a wall. The second rebuilt compare from the ground up: the original was shallow and hard to act on.

Live in Microsoft Foundry · Discover › Models › Compare
Shipped 01Region, promoted into the filter rail. An 11,000+ model catalog narrows by where a model can actually deploy, so region is a first-class filter instead of a wall you hit at deployment.
Shipped 02Compare, rebuilt from the ground up. Three models side by side with the best value in each row called out (the trophy), so trade-offs across quality, cost, and throughput read at a glance.

User Testing · Round 2

Validating the high-fidelity
prototype.

Round two tested one thing on the real prototype: could developers understand, navigate, and compare?

An earlier round on the wireframes had already shaped the layout; this final round put the high-fidelity prototype in front of developers across the Home, Discover, and Models pages, testing three things: whether they could understand the interface, navigate to the right region, model, or task, and compare options to find the best fit.

Three refinements came out of it: region and pricing moved to the top of each model card so they register before anyone opens a model, the compare tray became persistent with a running count, and the best value in each comparison row got called out so trade-offs read at a glance. Directional signals from a small, moderated round, enough to steer the design, not to claim statistical significance.

Results

The ideas that
are now shipping.

The two ideas I drove, region filtering and side-by-side compare, are now live for every Foundry developer.

The redesign reframed discovery around the decisions developers make: narrowing, evaluating, and committing, instead of scrolling an 11,000-row list and hoping. After the co-op, both ideas shipped into the live product that every Foundry developer now uses to choose a model: region sits in the filter bar, and there's a compare-model tray.

Live
Both ideas now shipped in the live product every Foundry developer uses to choose a model
11K→
An 11,000+ model catalog, filtered down to a focused, scannable set
~58%
faster search-to-select in a small, moderated round-two test, a directional signal rather than a significance-tested result

Task-success is the design metric; the reason it mattered to Foundry is the funnel. A developer's journey is browse › select › deploy, and two things quietly leaked users out of it: picking a model that turned out to be unavailable in their region (a failed deploy, and a reason to leave), and stalling in evaluation because comparing meant opening models one at a time. Region-in-filters removes the first leak before it happens; compare shortens the second. Both are aimed at the number the platform actually cares about, models successfully deployed, not just time on task.

Also during my Microsoft co-op: I ran the usability research for Foundry's Agent Builder, where 8 of 8 developers hit the same Severity-4 blocker, turning a deprioritized bug into a prioritized fix.

Read that study

Takeaways

What I took away.

01
Designing for emerging tech means continuous experimentation. There was no settled pattern for browsing 11,000 AI models, so we had to prototype our way to one.
02
Tight constraints sharpened the work. A fixed IA and no access to internal data forced focused decisions instead of a sprawling redesign.
03
Complexity can't always be eliminated, but it can be re-imagined. We couldn't shrink the catalog; we could change how people move through it.

Next Project

Nox+: Secure Digital Identity

View case study