Case study

Image data for visual search.

A computer-vision team needed two image datasets that models could actually learn from: 10,000 supermarket shelf images and 5,000 street-level storefront photographs, categorized on delivery.

15,000 - 10,000 images - 5,000 images

15,000 Total images delivered
10,000 images Supermarket shelf imagery
5,000 images Storefront imagery
Visual search image data visual: Speech transcription QA workflow with qualification, task tracking, and correction controls.
Measured outcomes Visual search image data
15,000 Total images delivered
10,000 images Supermarket shelf imagery
5,000 images Storefront imagery
Product type, brand, shelf configuration, storefront category, lighting Coverage dimensions
Categorized for direct pipeline ingestion Delivery state

Project overview

What landed, and what made it hard.

A computer-vision team needed two image datasets that models could actually learn from: 10,000 supermarket shelf images and 5,000 street-level storefront photographs, categorized on delivery.

Delivery snapshot

Visual search image data

Client
confidential AI product company
Service
Image data collection and categorization
Volume
15,000 images
Split
10,000 supermarket · 5,000 storefront
End use
Object detection and visual search training

Why this mattered

Outcome before process.

The two needs were related but distinct. Product detection from shelf imagery depends on seeing the same product category across different shelf configurations, packaging states, and store layouts. Storefront identification depends on seeing the same business type across different lighting, angles, and street conditions.

Both fail the same way. A dataset collected in one location under one lighting condition produces a model that works in that location under that lighting condition. Diversity is not a nice-to-have in visual training data; it is the difference between a model that generalizes and one that memorizes.

MoniSa handled the work under ISO 9001:2015 for process control and ISO 27001:2022 for information handling. Categorization labels required linguistic consistency, which the quality process enforced; ISO 17100:2015 is scoped to translation and is not claimed for this work.

The buyer-side requirement was that images arrive ready to ingest. A folder of unsorted photographs is raw material, not a dataset. Category structure had to match how the training pipeline expected to read it.

The problem to solve

Why the work was difficult, and what MoniSa changed in-flight.

Real-world image collection has a coverage problem that is easy to underestimate at brief stage. A buyer asks for 15,000 images. What determines model quality is not the count but the spread: how many distinct product categories, how many shelf configurations, how many lighting conditions, how many business types.

The challenge

The problem to solve

Supermarket imagery carries its own difficulty. Shelves are dense, products overlap, packaging reflects light, and the same product appears in different facings and stock states. A collection that captures only full, well-lit shelves trains a model that fails on the half-empty end-of-day shelf it will actually meet.

Storefront photography adds environmental variance that cannot be controlled. Daylight changes through the day. Signage sits at different heights and angles. Some businesses have glass frontage that reflects the street back at the camera.

The categorization layer is where datasets usually degrade. If category labels drift between collectors — one person filing a category one way and another filing it differently — the training pipeline inherits that inconsistency and the model learns the inconsistency along with the objects.

There is also a delivery-shape question buyers should ask early. Images organized by collection date are organized for the collector. Images organized by category are organized for the model. Those are different deliverables and only one of them is directly ingestible.

For a procurement team, the checks that matter here are specific: how many distinct categories, what environmental variance was deliberately captured, who defined the category taxonomy, and what the folder structure looks like on delivery.

Operating response

What MoniSa changed

MoniSa collected across multiple locations rather than concentrating the work where it was easiest to run. Location spread was the mechanism for the environmental diversity the training pipeline needed.

  • Location spread Collection ran across multiple locations so environmental and configuration variance was built into the dataset rather than corrected for afterwards.
  • Category-first structure Images were organized by category on delivery, matching how the training pipeline reads data rather than how the collection happened to run.
  • Deliberate condition variance Lighting, shelf configuration, and business type were treated as coverage dimensions, not incidental properties of wherever the camera happened to be.
  • Separate datasets, separate taxonomies Shelf imagery and storefront imagery kept distinct category structures because they train different model behaviours.

Results

Measured outcomes from this engagement.

15,000 images were delivered across the two datasets, categorized for direct ingestion into the object detection and visual search training pipelines.

Total images delivered15,000
Supermarket shelf imagery10,000 images
Storefront imagery5,000 images
Coverage dimensionsProduct type, brand, shelf configuration, storefront category, lighting
Delivery stateCategorized for direct pipeline ingestion

Selection logic

What protected the result.

The selection came down to whether MoniSa could source and review the work at standard, and whether that would hold across the full run.

Why the fit was real

Why the fit was real

The work needed multi-location collection capacity plus a categorization discipline that produced an ingestible dataset rather than a folder of photographs.

What decided the result

What decided the result

Deliberate variance across lighting, configuration, and category did more for training value than raw image count.

What buyers can reuse

What buyers can reuse

  • Judge an image dataset by its coverage dimensions, not its image count. 15,000 images from one location under one lighting condition is a smaller dataset than the number suggests.
  • Ask who defines the category taxonomy and when. A taxonomy agreed after collection produces re-sorting work the buyer pays for twice.
  • Two datasets serving two model behaviours need two structures. Merging them makes the headline number bigger and the data harder to use.
  • Environmental variance has to be collected deliberately. It cannot be added later, and a clean dataset trains a model that fails on messy inputs.
  • A useful image-collection brief names the categories, the required variance per category, the delivery folder structure, and the ingestion format before collection starts.
  • Ask for the category structure as a sample before volume runs. It is far cheaper to correct a taxonomy at 100 images than at 15,000.
  • Collection capacity and dataset usability are separate capabilities. Verify both.

Continue from this proof

Useful comparisons for the same problem.

Use these links to compare the case with the matching service, buyer guide, and language coverage.

Languages named

Examples referenced in the engagement.

  • Retail product imagery
  • Storefront photography
  • Category taxonomy
  • Object detection training data

case evidence

Nearest proof pattern.

These related cases keep the next click close to the same kind of work.

AI data servicesA single voice data project became a recurring relationship, including rare-language work.

Voice data, 500 hours

The challenge. A client needed ~500 hours of US English voice data with diversity and privacy requirements, inside a fixed budget.

What we did. MoniSa collected inside the client's own app environment with privacy controls, two-layer QC, and terms that protected speaker diversity.

The result. ~500 hours at reviewed quality satisfaction, extended by the client into follow-on rare-language engagements.

Open full case
AI data services178 annotated hours delivered with 11 of 20 recruits removed before they touched production data.

Bengali pilot, screened

Problem. An AI company needed a Bangladeshi Bengali annotation pilot on a fixed timeline, where dialect and annotation aptitude are separate requirements.

Action. MoniSa over-recruited, screened aptitude separately from fluency, and removed those who did not meet standard before production.

Result. 9 of 20 cleared screening; 178 hours delivered as paid production work with the funnel reported in full.

Open full case
AI data services~300 hours of voice bot conversation transcribed with disfluencies preserved for training value.

Voice-bot transcription

Problem. A partner needed human-side conversational transcription across six languages, where cleaning the transcript would destroy the training signal.

Action. MoniSa held a verbatim convention across six per-language standards and kept the two Spanish variants operationally separate.

Result. ~50 hours per language delivered, with false starts and corrections intact.

Open full case

Buyer questions

Ask the questions weak vendors avoid.

Short answers for buyers checking fit, coverage, quality method, and next-step readiness.

What was delivered on this engagement?

Total images delivered: 15,000. Supermarket shelf imagery: 10,000 images. Storefront imagery: 5,000 images

What control kept the work stable?

Deliberate variance across lighting, configuration, and category did more for training value than raw image count.

Where should similar work go next?

Use AI data services for the delivery model, AI data annotation buyer guide for buyer-side evaluation, and the contact page for a scoped brief.

Similar brief

Send the constraint behind the metric.

A useful follow-up to a case study names the language mix, review model, deadline, and what proof your buyer team needs before approval.

Production-ready brief

01Closest matching challenge from this case02Language pair, dialect, and script coverage03Volume, cadence, or hours to deliver04Reviewer model and acceptance criteria05Security or platform constraints06Proof needed for stakeholder approval

Capability at a glance

The answers most briefs open by asking for.

Buyers rarely start with who we are. They start with a list of fields to fill. Here are ours, so the first email can be about the work instead.

Languages and locales
300+ languages and 4,500+ dialects, quoted per locale rather than per language — because the dialect decides whether a dataset is usable, whether a market accepts a release, and which specialist the work goes to.
Specialist network
110,000+ verified language specialists — linguists, annotators, and reviewers — plus voice talent and subtitlers, matched to the language, domain and task before assignment.
Capacity and mobilisation
Named availability confirmed per pair before scoping. Coverage is reported as staffed today or needing a recruitment window, in writing, before a launch date or release window is agreed.
Sourcing constraints
Specialists can be sourced against geographic, residency, locale and demographic requirements — including native-only, in-country, and speaker-diversity quotas where a data programme demands them.
Deliverables and specs
Work is delivered to the receiving specification: structured formats and schemas for data and annotation work, and timed-text, audio and platform conformance for media — subtitle reading speed, line limits, cue timing, channel and sample-rate requirements included.
Comparable work
62 documented case studies stating the scope, the constraint that made it difficult, and the measured result — across AI data programmes, partner overflow, and media releases. 2,000+ AI projects delivered and 1,000+ brands served since 2015.
Certifications
ISO 9001:2015 quality management, ISO 27001:2022 information security, and ISO 17100 translation services — scoped to translation specifically, and stated that way rather than implied across every line.
Commercial basis
Quoted in the unit the work is measured in — per word, per audio hour, per approved hour, per finished minute, per batch, per item — with what the unit includes stated alongside it, whether the quote is for you or for a client you quote onward.

Need this against your own template? Convert your scope between units and check the deadline, then send the brief with your language list, content type, volume and deadline, and the acceptance criteria you will judge the output against — those four decide feasibility, and the reply addresses them directly.

Scope a project Call