E-commerce · Final-year project · 2025 · Not live

ShopAgent

A voice shopping assistant for Android: say “Hey ShopAgent”, name what’s on screen, and it finds where to buy it.

0:29
per detection, down from 11s
0.4s
faster on cached repeat requests
47.8×
less GPU memory for Grounding DINO
61.8%
of GPU memory holds both models
1.7 GB

The problem

Buying something you see in a post meant screenshotting it, cropping it and running a reverse image search by hand. Open-source models could cut an item out of an image from a description, but the starting demo took 11 seconds per image.

Who it’s for
Android users who want to buy what they see in other apps, without leaving the app they are in.
My role
Sole AI/ML engineer
Team
Solo
Timeline
1 month

Stack

Models

  • Grounding DINO
  • SAM 2.1

App

  • Flutter
  • Kotlin
  • Picovoice
  • Android speech

Backend

  • FastAPI
  • PyTorch
  • Hugging Face
  • spaCy
  • WebSockets

Search

  • Google Lens
  • SerpAPI
  • Cloudinary

How it works

Phone

  • Wake word
  • Screen capture
  • Speech to text

API

  • Detect endpoint
  • WebSocket

Vision

  • Prompt parser
  • Grounding DINO
  • SAM 2.1
  • Mask render

Search

  • Item crop
  • Cloudinary upload
  • Google Lens

App

  • Results sheet
  • Store page
  1. 01 “Hey ShopAgent” from any app

    Picovoice hears the wake word on the phone, and a background service grabs the screen before ShopAgent comes forward.

  2. 02 The request becomes a prompt

    Android speech-to-text transcribes the request; on the server, spaCy cuts it down to the item, such as “green tank top.”

  3. 03 Two models find the item

    Grounding DINO boxes the object the words describe, and SAM 2.1 turns the best-scoring box into a pixel mask.

  4. 04 The cut-out arrives first

    The mask is rendered in magenta and pushed over the WebSocket while the product search is still running.

  5. 05 Google Lens finds the shops

    The top box is cropped, uploaded to Cloudinary for a public URL, then searched with Google Lens through SerpAPI.

  6. 06 Matches fill the results sheet

    The matches come back in the HTTP response, and each card opens the store’s page.

Decisions

01

Grab the screen first

Chose Screen capture inside the wake-word service over opening ShopAgent first and capturing after.

The item the user means is in the app they were using; once ShopAgent comes to the front, that screen is gone.

Trade-off: It needs a foreground service plus screen-capture and overlay permissions, which Android makes users grant one by one.

02

Load the models once

Chose Loading both models once at server start-up over loading them on every request, like the demo script.

Loading the two models took 10.9s of the demo’s 11s; running them takes about 0.4s.

Trade-off: The models hold about 1.7 GB of GPU memory even when idle, so they unload after 30 idle minutes and the next request reloads them.

03

Stream the cut-out early

Chose Sending the mask early through a WebSocket over returning everything in one HTTP response.

The user sees the item cut out as soon as SAM 2.1 finishes, instead of waiting for the upload and web search.

Trade-off: The app has to open the socket before it posts, and the two halves are matched on a request ID.

What I’d do next

  • I’d give every request its own files. Each upload is saved to the same path, so two people searching at once would overwrite each other’s images.
  • I’d build an accuracy set before tuning speed. The repo times every model and the cache, but never checks that the right item was found.
  • I’d load the spaCy parser once at start-up, like the vision models. It still reloads on every request.

Results

  • Detection time per request

    Before: 11sAfter: 0.4s

  • Model loading per request

    Before: 10.9sAfter: 0s

  • Repeat request speed, cached

    Before: 1×After: 47.8×

  • Grounding DINO speed

    Before: 1×After: 1.4×

More work