Android Gemini Patterns

Early Gemini was a capability in search of an interface. This was me spending a few weeks figuring out what that interface might actually feel like not in theory, but in your hand, on your phone.

  • Role: Interaction Designer / UX Concept Lead
  • Company: BUCK (for Google)
  • Timeline: November – December 2023
  • Platform: Android (mobile)
  • Tools: Figma, hand sketching, concept diagramming
  • Status: Delivered as internal concept exploration for Google's early Gemini product team. Several interaction patterns and mechanics from this work have since surfaced in Gemini features across Android OS releases.
  • Team: Freelance engagement within a BUCK concept team: Creative Director, two motion designers, and a rotating group of additional designers. I was the dedicated interaction designer on the project, brought in to develop as many of these concepts through to defined interaction patterns as possible.

The brief #

In late 2023, Google had early Gemini capabilities but no clear native UX model for Android. The question wasn't "can AI do this," it was "how does a person reach for this on a phone." BUCK was engaged to pressure-test that question through real interaction patterns: not a spec, not a feature list, a working exploration of what the paradigm could actually be.

Existing AI interfaces in late 2023 were mostly chat boxes. That's fine at a desk, but on a phone it's wrong. The phone is physical, gestural, contextual. If Gemini was going to live on Android it needed to feel like Android.

My role #

BUCK brought me in as a freelance interaction designer with a specific mandate: Google had dropped a significant body of early thinking on the agency and needed someone to work through it systematically. Figuring out which directions had real interaction legs, and developing as many of those as possible into defined, demonstrable patterns. I worked closely with a creative director and two motion designers, with other designers pulling in and out across the engagement.

My contribution was the interaction design and concept development throughline. Taking loose scenarios and frameworks and pressure-testing them into concrete gesture-based UX patterns. I owned the exploration process: framing the right questions, generating concepts across the scenario space, iterating the interaction models, and distilling the output into flows Google could actually use as directional material.

Drive-to-output #

The framing question that drove everything: what is the human objective, what does the existing process cost them, and how does AI change the gesture? Starting from that, I mapped scenarios where AI would actually change what someone does, not just speed up something they already do.

The first week's scenarios covered things like shopping for a winter jacket (the whole research-compare-decide loop collapsed into a weighted query you could tune on a slider), deck building from a folder of class notes (Holistic context menu → select content → describe the goal by voice → reviewed deck), message sorting (shake to break the reverse-chronological thread apart into meaningful groupings), and camera as a universal remote (point at a device, the phone figures out how to connect).

Holistic context menu, resting state and broader folder actions

Describe the goal, hold to add direction, listening state, output deck

The "Drive to Deck" flow was one of the cleaner end-to-end ideas: you're in your Drive folder, you ask it to make a presentation, it surfaces a Gemini panel where you can either tap an action or hold and speak a goal "I want to build a presentation of vacation photos from my Mexico trip with historical data from all the places I saw", and it listens, surfaces the key intent back to you highlighted as a chip, and then generates the deck.

After that initial output there are a few different refinement paths: Tune the goal (sliders for detail vs length, images vs text, or theme selection), Shake to Refine (breaks the deck apart by content type so you can reorganize before squeezing it back together), or just take the optional outputs, share deck with notes, or present with audio recording and transcription.

Tune the goal, color palette, typeface, emphasis tuning Shake to Refine, break out into analyzation & grouping, user edits, squeeze back

The video editing version of this ran the same pattern: Drive to Video, Shake to Refine, Rotate to Tune (landscape breaks into a drag-and-drop timeline), Zoom on Text (expand transcript for collaborative script editing), Hold & Drag to Tune (XY graph for a clip's emphasis: color vs audio correction, video vs motion graphics), Update Seed Content.

Drive to Video, analyze & group key scenes, faces, transcription Shake to Refine (video), break back up, remembers analyzation chunks for collaborative editing
Rotate to Tune, landscape timeline editing, drag & drop Hold & drag to Tune, XY graph emphasis dial on a clip
Zoom on text, expand transcript, collaborative script editing, pinch to see refined video Update seed content, new clips emphasized, missing content deprioritized but remembered

The camera layer #

The second week went deep into the camera layer. The hypothesis was simple: if Gemini is going to live on Android, it needs to recognize subjects in photos the way a person would and then give you a way to act on them without typing.

The earliest concept frames had both you and your dog rigged simultaneously. Two subjects, each with their own pink dot and white skeleton lines extending to joint points. The system sees everything in the frame, not just the person you tapped on.

Early icon radial, three unlabeled bubbles: Image, Video, Text icons

Refined bubble radial, Pose / Emotion / Background / Style, Gemini spark center

Full labeled action state, green/blue/yellow pills, color-coded by action type

Iteration 1, Icon radial. Tap, and three unlabeled bubbles bloom out from the spark: Image, Video, Text (icon-only). No labels yet, just the shape of the system.

Iteration 2, Labeled icon radial. Same structure, with labels. "Video / Image / Text" in the bubbles. A profile thumbnail of you appears at the bottom, the system knows who it's looking at and surfaces your identity as context.

Iteration 3, Full action pills. The system opens up into a full labeled state: green pills for subject-level actions (Change Emotion, Repose Person, Effects Filter), blue for gallery actions (Show me more photos, Add someone else), yellow for social (Chat with Jonny). Color-coded by intent, the pill shape is distinct from the earlier bubble approach.

One alternative explored at this stage was zone mapping instead of a radial menu, the photo itself becomes the selector. The face region highlights yellow (tap for Emotion), the body highlights teal (tap for Pose), the background pinks out (Background), a blue diamond appears at the top (Style). You pick by tapping the zone on the actual subject rather than a separate UI element. Different mental model entirely.

The iteration that stuck was a tighter bubble radial, Pose / Emotion / Background / Style centered on the Gemini spark with a clean X to dismiss. Much less noise than the full pill state.

Zone-mapped selector, body regions color-coded as tappable action zones: Emotion/Pose/Background/Style Wireframe rig, body pose editing with draggable joint handles Green landmark rig, accepted pose, facial feature points, ready to apply

From here, each action branches into its own selector:

Selecting Pose drops a white wireframe rig onto the person, skeleton lines, joint dots, a grid oval over the face. You drag the joints to repose. Once the rig is confirmed, it switches to green facial landmark dots, feature points at eye corners, nose bridge, jaw, skeleton extending to shoulders. Accepted, ready to render.

Selecting Emotion opens an emoji ring, ~20 emoji floating in a circle around a green center dot, overlaid on the face. You spin the ring to land on the target expression. Not a dropdown, not a text field. You pick 😎 or 😂 and that becomes the generation target.

The flip side of that is selecting from your own face history, "Show me more photos" surfaces a ring of your actual face crops from the camera roll, each one a different hat, expression, lighting condition. Same green center dot, but instead of emoji you're picking a reference photo to move toward. Two ways to say the same thing: one abstract (emoji), one literal (your own face).

Color picker, drag the handle across the full spectrum to set the grade Emoji ring, spin to select target emotion, overlaid on the face Show me more, face thumbnails from camera roll arranged as a ring around the green center dot

Selecting Effects Filter opens a full-spectrum color blob with a white circle handle. Drag it across the spectrum to pick the mood/grade. No preset names, no slider labels. Direct.

Selecting Show me more photos expands into a thumbnail strip of every photo of you from your camera roll, different hats, lighting, expressions, plus emoji bubbles and large "Show me more" / "Add someone" CTA buttons at the bottom. The AI is showing you a selection to move toward rather than asking you to describe what you want.

The earlier hand-annotated sketch that framed all of this, yellow callouts on a selfie: Expressions, Back Ground, Pose Character, Style Transfer (Manga Comic), was the clearest version of the concept. Sometimes the sketch is the spec.

Show me more, camera roll thumbnail strip, emoji bubbles, Show me more / Add someone buttons

Working sketch, USER SELECTS POSE OWNACTZ, yellow and blue rig markers on face close-up

Hand-annotated selfie sketch, face rig concept, drawn directly on the photo

There was also a parallel concept: Gemini as a floating presence object, a 3D geodesic sphere with colored action pills radiating from it. Less of a tap-and-reveal thing, more of a persistent ambient entity that surfaces relevant actions based on context. Different direction than the tap-on-a-subject model, but worth keeping in the picture.

Gemini globe, 3D geodesic sphere with color-coded pill actions radiating from it

The Dots & Handles abstraction came out of all of this, the same rig-point idea applied beyond photos. Face points for animation, pitch handles on audio to swap a voice actor, anchor points on a room photo to shift the aesthetic. The concept was the thread that ran through everything: AI shouldn't just generate output. It should give you something to grab onto after it does. Every polished interaction on this project was trying to get to that feeling the moment where the phone hands the result back to you and it's still malleable.

Dots & Handles, 3-screen AI rig concept, applied across media types

The Supercut persona screen was a different direction in the same session, Gemini surfacing a collage built from your phone data: faces of people you've photographed, a map of your home area, books, cities, saved rooms. The argument being that an AI that actually knows you shouldn't need a form. It should just show you what it's built.

Supercut persona screen, Gemini contextual collage from phone data: people, places, books, cities

Whats That & Refine #

Another big concept was Whats That, a style/generation UI for when you already have a result but want to move within its aesthetic space without typing a new prompt. A spectrum bar runs across the bottom of the screen. Drag the handle and you slide through aesthetic interpretations: for a room it goes Hollywood → Midcentury → Industrial → Art Deco → Cottage. For an outfit: Gorpcore → Energetic Techwear → Miami Y2K. The whole language of style collapsed into a single draggable gesture.

Whats That, style spectrum slider, Cottage selected Whats That, slider moved, different aesthetic selected Whats That, slider at another position, full range visible

The companion to Whats That was the Refine concept, not about tuning a result, but tuning the AI itself. A radar/spider chart with six personality axes: friendly, simple, light, assertive, amused, formal. The shape you draw on the chart becomes the filter on how Gemini interprets and responds to everything. Not a settings panel, more like giving the model a personality dial you can reshape. The framing: instead of refining your results, you manipulate a set of simple parameters to adjust your entire search personality. The prompt field stays blank, you're not writing more text, you're moving handles.

Daily explorations #

Daily explorations covered scenarios that didn't reach polished flows but that held up as interaction concepts.

Shake to sort Messages reverse-chronological messages are the wrong model. Shake the list to break it apart by media type and conversational type, then sort out what's a fun convo vs a functional one.

Camera object recognition, Universal remote camera recognizes a device and reaches out to it, figures out how to connect and use it. Point at a speaker, your TV, a camera, a thermostat.

Pinch to summarize and navigate looking at a website or article, pinch to turn it into a bucket your AI can interact with. The polished version landed on a COP28 NYT article with dark pill buttons floating over the content: "Summarize this article," "Play a 2 minute summary," "Video recap of this," "Who is in this story?" surfaced without leaving the page or opening a new interface. The expanded state went further: "Play a 2 minute summary" is now active with an audio waveform player rendering inline, "Send this to Erika" appears as a yellow contact pill (Gemini sees who the article is relevant to based on your contacts), and "Can I trust this article?" becomes available alongside "Who is in this story?" A lot packed into a very small footprint.

Gemini over NYT article, action pills: Summarize, Play summary, Who is in this story? Article screen expanded, audio summary playing, Send to Erika contact pill, Can I trust this article? Refine, Gemini personality radar chart: friendly / simple / light / assertive / amused / formal

Influence #

This was concept work, not a production spec, so there's no shipping announcement to point to. But over the following year and a half, using Gemini features across Android, a number of the mechanics explored here have surfaced: gesture-driven AI actions on subjects, contextual action pills that radiate from a recognition point, refinement patterns that let you tune AI output rather than re-prompt. Whether these directions arrived at Google independently or through this work isn't something I can claim definitively. What I can say is the design space we were exploring in late 2023 was the right one.

Scope delivered: 7 interaction pattern families across 4 major scenario clusters developed from initial scenario mapping through polished interaction flows, Drive-to-output, camera-as-AI-interface, contextual web actions, AI personality tuning.

Reflection #

Looking back at it, this was one of those projects where the brief was intentionally loose because the territory was, and still is, new. The scenarios that resonated most were the ones where the AI was reducing the effort of a specific moment rather than adding a general capability. Whats That works because it collapses a search-and-iterate loop into a single draggable handle. Dots & Handles works because it makes AI output feel editable rather than final.

The stuff that didn't land as clearly was anything framed as "here's an existing feature with AI added." That always came out a little forced. The stronger work started from behavior, not feature.

There was more to do with the physical gesture language: hold, fling, pinch, shake, rotate. Android already has a rich physical vocabulary and this was starting to use it, but didn't fully commit. The rotate to landscape → timeline edit flow was probably the clearest version of that. That's the direction I'd want to push further.

Working in this space early, before the production constraints settled, before the platform conventions hardened, is a particular kind of design opportunity. You're not solving a problem yet. You're deciding what problems are worth solving.