[Case 03]

From ~6 hours to minutes: In-app incident communication for Credit

I turned a request for one "maintenance screen" into a scope-based incident system. Built with Engineering using SDUI and feature flags, so comms could go live in minutes across mobile and web, with no release dependency.

Scale: 120 incidents (16 critical), ~36k customers affected, Sep '24 to Mar '25

Impact: ~36k customers impacted
/ 120 incidents (Sep–Mar)

[Industry]

Fintech

[My Role]

Product Designer

[Platform]

Mobile + Web

Context

Credit is a high-trust domain: contracting, payments, renegotiation. When something breaks there, customers read it as risk to their money.


Between Sep '24 and Mar '25, Credit had 120 incidents (16 critical) affecting ~36k customers (Looker Studio). Downtime was expensive on its own. The bigger cost was ambiguity: people couldn't tell what was still safe to use, so they retried in loops and escalated to Support.

Problem

The request that came to me was "design a maintenance screen." But credit had too many moving parts for one screen to cover:

  • Three products that could fail independently (Giro Fácil, Capital de Giro, Renegotiation)

  • Five entry points where users could hit the failure (home grid, Credit hub, top card, inside a flow, deep link)

  • Three failure levels: credit-wide, product-level, action-level

Incident comms were slow and scattered. Customers hit a generic error or no guidance at all, retried, and ended up in Support. And because most of those surfaces shipped with the app, the message often arrived hours after the team already knew.

Before

Different entry points collapsed

into the same dead end

After

My contribution
  • Led end-to-end UX/UI for incident communication across mobile and web (Credit)

  • Turned a one-screen request into a reusable incident system

  • Mapped 50+ scenarios (scope × entry point × eligibility) to catch what happens outside the happy path

  • Shipped with Engineering using SDUI and feature flags, enabling activation in minutes

Process

I anchored the system on two levers:

  1. Scope (what's affected): credit-wide, product, or action

  2. Entry point (how the user got there), which sets the right friction level, from a quick check to a deep dive

In Support check-ins, the same questions kept coming up:

Is this temporary or did I lose access? What exactly is affected? How fresh is this info?

So I designed every surface to answer those, and used "Last updated" in place of ETAs we couldn't keep accurate.

Benchmarks: Layered comms is the norm

Benchmarks: Layered comms is the norm

Benchmarks: Layered comms is the norm

Internal: We had pieces, not a system

Internal: We had pieces, not a system

Internal: We had pieces, not a system

Support: Users ask the same 3 things

Support: Users ask the same 3 things

Support: Users ask the same 3 things

Solution

1) Scope: what failed

Credit-wide · Product-level · Action-level

Scope decides what gets blocked and what stays usable. It also sets what the details page has to explain.

2) Entry point: how the user arrived

Home grid · Credit hub · Top card · Inside-flow · Deep link

Entry point sets the friction. Someone just checking gets a quick status. The full explanation only shows up when they need it.


From there I shipped four reusable surfaces:

  • Quick status (bottom sheet) for people who are only checking

  • In-context message at the point of failure, so the rest of the product stays usable

  • Entry intercept when the whole area is down

  • Canonical details page carrying scope, last updated, and next steps

Why no ETA: we almost never had a reliable "back at X time," and keeping a countdown accurate in the UI would have been expensive and easy to get wrong. We used "as soon as possible" plus a "last updated" timestamp, so the reassurance came from information we could actually keep fresh.

Impact

Speed: Engineering reported that time to publish in-app incident comms dropped from ~6 hours to minutes (engineering metric: comms publish time).

Leverage: The pattern became a reusable SDUI incident system other teams could pick up for their own domains.

What we set up to measure (post-launch): incident UI reach and user actions, support contact rate among exposed users, and publish speed during incidents.

Learnings & next steps

Reliability communication needs the same care as any core product flow. Scope and entry point were what kept it consistent at scale without blocking more than necessary.

Two things I'd fix with more time:

  • Operational autonomy: let Support and Comms update status and "last updated" through a lightweight admin control, so it doesn't sit on Engineering.

  • Company-wide consistency: test the pattern in domains outside Credit to see if it holds for a broader set of error scenarios.

Select this text to see the highlight effect