AI-Powered Automatic Content Tagging in Mobile Apps
A typical scenario: a user uploads a photo to an app, but to add a tag, they must manually choose from hundreds of options or type text. Context and time are lost. We implement automatic AI-based tagging — images, text, and video get labels without human intervention. Our track record: over 20 projects, from marketplaces to social networks, with custom taxonomies. The solution applies to any domain: medicine, real estate, retail, education. On-device models achieve 92% accuracy, server models 98% with a properly tuned confidence score. We guarantee 98% accuracy threshold adjustment and have over 10 years of experience in AI and mobile development.
For example, an online clothing store: we trained a model to recognize 150 categories with 91% accuracy. Now every uploaded item automatically receives tags like “Dress”, “Cotton”, “Summer”. This cut moderation time by 4x and saved roughly $4.5k–6.5k per year in manual labeling. Our clients typically save $10,000–$50,000 per year depending on content volume.
Apple Core ML documentation recommends transfer learning for creating compact models with 20+ examples per category.
How AI Tags Images in iOS and Android
The standard stack is VNClassifyImageRequest (iOS) + ImageLabeler (Android). They return generic labels like “Food”, “Sky”, “Cat”. For business needs, you need a custom taxonomy: not “Clothing”, but “Leather jacket”, “Floral dress”. We train a custom model using CreateML (iOS) or TensorFlow Lite (Android). Below is an example of training a custom model under iOS.
// Training via CreateML (run on Mac, not on device) import CreateML let trainingData = MLImageClassifier.DataSource.labeledDirectories( at: URL(fileURLWithPath: "/training_data") // Structure: /training_data/jacket/, /training_data/shoes/, /training_data/bag/ ) var params = MLImageClassifier.ModelParameters() params.maxIterations = 25 params.validationData = .split(strategy: .automatic) params.featureExtractor = .scenePrint(revision: 2) // Transfer learning from Apple let model = try MLImageClassifier(trainingData: trainingData, parameters: params) try model.write(to: URL(fileURLWithPath: "/model.mlmodel"), metadata: nil) 20–50 examples per category, 15–30 minutes of training on a MacBook Pro M2 — you get a compact model. Core ML Model Deployment allows updating it without publishing a new App Store version.
Why Hierarchical Tags Speed Up Search
A flat list of tags is chaos. A hierarchy like “Food → Italian cuisine → Pasta” gives structured search and filters. We implement this via trees:
// Android: TagTree data class Tag( val id: String, val name: String, val parentId: String?, val synonyms: List<String> = emptyList() ) // When tagging: if tag "Pasta" is assigned, automatically add parent tags fun expandWithParents(tagId: String, tagTree: Map<String, Tag>): Set<String> { val result = mutableSetOf(tagId) var current = tagTree[tagId] while (current?.parentId != null) { current = tagTree[current.parentId] current?.let { result.add(it.id) } } return result } For storage we use a separate table with a source field (auto, user, admin). Auto-tags are visible only in search, user tags in the UI.
On-Device vs Server Tagging Comparison
| Characteristic | On-device (CreateML / TensorFlow Lite) | Server-side (OpenAI / Claude) |
|---|---|---|
| Latency | Instant (5–50 ms) | 0.5–2 s |
| Offline mode | Yes | No |
| Privacy | Data never leaves device | Data goes to server |
| Accuracy | 85–92% on narrow taxonomy | 95–98% on complex requests |
| Cost | Free (device compute resources) | Pay per API request |
We combine both: basic tags are set on the device, for complex cases we send a request to the server. On-device tagging is 3–10x faster than server with similar accuracy for typical categories.
How does confidence score work?
Each model returns a probability for each category from 0 to 1. We set a threshold (usually 0.7–0.9) — tags below the threshold are dropped. An administrator can review and correct auto-tags. A/B testing different thresholds helps find the optimal balance between precision and recall.How to Improve Tag Accuracy
If accuracy is below expectations, increase the training set to 100+ examples per category or use a server-side model for difficult cases. Regular retraining on new data keeps the taxonomy up-to-date. We recommend retraining the model monthly as new content types appear.
Tagging Text and Video
Text posts — NLP classification on-device via the Natural Language Framework or on the server. Prompt: "Determine 3–5 tags from the list: ...". JSON response is parsed on the client.
Video — key frame analysis:
func tagVideo(at url: URL) async throws -> Set<String> { let asset = AVURLAsset(url: url) let duration = asset.duration.seconds let generator = AVAssetImageGenerator(asset: asset) generator.maximumSize = CGSize(width: 224, height: 224) var allTags = Set<String>() var time = 0.0 while time < duration { let cgImage = try generator.copyCGImage(at: CMTime(seconds: time, preferredTimescale: 600), actualTime: nil) let frameTags = try await classifyImage(cgImage) allTags.formUnion(frameTags) time += 3.0 // every 3 seconds } return allTags } For long videos we use background tasks (BackgroundFetch on iOS, WorkManager on Android) or send to the backend.
Implementation Steps
- Content Audit & Requirements: Analyze current content volume and types, define business goals.
- Taxonomy Design: Create flat or hierarchical tag structure with stakeholder input.
- Data Collection & Annotation: Gather minimum 20 examples per category, annotate manually.
- Model Training: Use CreateML (iOS) or TensorFlow Lite (Android) with transfer learning; validate accuracy.
- Integration: Embed SDK into app, connect on-device and server models, implement search filters.
- Testing & Deployment: A/B test thresholds, deploy to production, monitor performance.
Deliverables
- Audit report and taxonomy design document
- Annotated training dataset
- Trained model (.mlmodel / .tflite) and source code
- Integrated iOS (Swift) and Android (Kotlin) SDK with hierarchy support
- On-device + server architecture (Firebase, Supabase, or your backend)
- Accuracy testing results and threshold recommendations
- Complete API and taxonomy documentation
- Team training (2 workshops)
- One month post-release technical support with access to source code and model artifacts
Implementation Timeframes
| Stage | Duration |
|---|---|
| On-device image tagging (ready-made models) | 3–5 days |
| Custom taxonomy + domain-specific training | 1–2 weeks |
| Text + video tagging + hierarchy | 2–4 weeks |
| Full cycle (analytics → design → test → deploy) | 3 to 8 weeks |
Pricing is determined individually. For an accurate estimate, send us your project description — we will prepare a commercial proposal within 1–2 days.







