Condé Nast’s editorial teams had no fast way to do multimodal video discovery. They were spending an average of 250 minutes per content discovery task, manually scrubbing through a library of more than 140,000 videos. They relied on titles and descriptions to find relevant clips. In a media environment where speed-to-market directly determines revenue capture, this process created measurable operational drag across brands such as Vogue, GQ, Vanity Fair, and Wired.
The core problem was structural: Existing search tools can’t look inside video content. Teams depended on institutional knowledge to locate assets, creating single points of failure when specific individuals were unavailable. Meanwhile, underutilized content sat in the archive undiscoverable because no keyword in a title or description connected it to the queries editors were actually running.
To solve this, Condé Nast partnered with the AWS Generative AI Innovation Center (GenAIIC) to build an AI-powered multimodal video discovery solution. Built on Amazon Bedrock and Amazon OpenSearch Service, the solution runs intent-based semantic search across video transcripts, visual elements, and audio. The team selected the TwelveLabs Marengo embedding model for its native ability to jointly encode visual, audio, and transcript signals. Marengo powers all five capabilities described in the following section. This reduced discovery time from 250 minutes to under 2 minutes per task.
In this post, we describe the architecture, explain why we chose specific technology choices, and share the business outcomes the solution delivered.
Why semantic search (and why this architecture)
When the team scoped the problem, two constraints shaped the solution design.
First, keyword search was fundamentally insufficient. Editorial teams don’t search for “yoga_tutorial_march_2024.mp4.” Instead, they search for “beginner yoga content with calming backgrounds” or “behind-the-scenes fashion week moments.” The search layer needed to understand intent, not match strings. This pointed directly to vector embeddings that capture semantic meaning across modalities (visual, audio, and transcript). Human-authored metadata alone could not provide that level of understanding.
Second, the video library was large enough (over 140,000 videos) that any solution needed to separate the expensive, compute-heavy work of generating embeddings from the low-latency work of serving search results. A monolithic architecture would force a tradeoff between ingestion throughput and query responsiveness. Decoupling them meant each plane could scale, fail, and evolve independently. This decision proved essential during backfill processing.
These two constraints led to the solution’s core design: multimodal vector embeddings generated by the TwelveLabs Marengo model on Amazon Bedrock, indexed in Amazon OpenSearch Service, and served through a purpose-built query tier. This tier converts natural language into vector searches and returns precise timestamps. We accessed Marengo through Amazon Bedrock because Bedrock offers a range of foundation models (FMs) through a single API. It also applies AWS governance controls, including AWS Identity and Access Management (IAM) for access, Amazon Virtual Private Cloud (Amazon VPC) for network isolation, and AWS CloudTrail for auditability. That combination let the team adopt a specialized embedding model without building or operating its own model-serving infrastructure. For the vector index, Amazon OpenSearch Service provides managed k-nearest neighbor (k-NN) search with multi-AZ replication and metadata filtering for hybrid queries. This let the team run low-latency similarity search across more than 140,000 videos without managing the underlying search cluster.
Solution capabilities
The solution gives editorial teams a fundamentally different way to interact with their video library. The following capabilities are available through the solution:
Intent-based search: Users describe what they need in natural language (for example, “beginner yoga content” or “celebrity interview about sustainability”) and receive contextually relevant clips with precise timestamps.
Multimodal understanding: Search spans video transcripts, visual elements, and audio simultaneously. A clip is discoverable by what’s said, what’s shown, or what’s heard.
Image-based queries: Users can upload a reference image to find visually similar content across the archive.
Typo tolerance and intent resolution: The search layer handles imprecise queries and focuses on meaning rather than exact keyword matches.
Timestamp precision: Results pinpoint exact moments within videos, alleviating the need to watch entire clips.
The following diagram illustrates the high-level architecture, showing the ingestion and indexing plane and the query and serving plane.
Figure 1: High-level architecture of the multimodal video discovery solution, showing the ingestion and indexing plane and the query and serving plane.
Architecture walkthrough
The solution consists of two decoupled planes: an asynchronous ingestion pipeline that makes videos searchable, and a synchronous serving tier that handles user queries.
Ingestion and indexing plane
When a new video is uploaded, the following sequence executes:
Upload: The video lands in an Amazon Simple Storage Service (Amazon S3) bucket that holds source video (Video Outputs in the diagram).
Orchestration: A new upload raises an event (Video Search Events in the diagram). That event pushes the video’s metadata into the ingestion process and coordinates the pipeline.
Preprocessing: The ingestion service validates the video and extracts metadata (format, duration, resolution).
Chunking: The same ingestion service, running on Amazon Elastic Container Service (Amazon ECS) with AWS Fargate, splits the video into segments. An Auto Scaling group scales it horizontally across multiple Availability Zones to process chunks in parallel.
Embedding generation: Each chunk is processed through the TwelveLabs Marengo embedding model on Amazon Bedrock using asynchronous invocation. Marengo generates multimodal vector embeddings capturing the semantic content of each segment across visual, audio, and transcript dimensions.
Indexing: Generated embeddings are written to an Amazon S3 bucket for durability (Video Embeddings in the diagram) and indexed into an Amazon OpenSearch Service cluster for vector similarity search.
The pipeline is event-driven and orchestrated end to end. This gives the team per-step retry logic, parallel processing across chunks, and full traceability for debugging. These were capabilities the team relied on heavily during the initial 140,000-video backfill.
Query and serving plane
When a user performs a search, the following sequence executes:
Request routing: The request enters the VPC through an internet gateway and reaches the external Application Load Balancer (ALB).
Backend processing: The external ALB distributes requests to the frontend service on Amazon ECS with AWS Fargate. The frontend service calls the search service through an internal ALB. Auto Scaling groups scale both services across multiple Availability Zones.
Vector search: The search service converts the user’s query into a vector embedding and performs a k-NN similarity search against the Amazon OpenSearch Service index.
Metadata enrichment: The search service reads video metadata (titles, thumbnails, references) from Amazon DocumentDB (with MongoDB compatibility) to enrich results.
Response: Combined results return to the user with relevant clips and precise timestamps.
High availability (HA) design
The solution is multi-AZ throughout. The following components are deployed for resilience:
Amazon OpenSearch Service: Synchronous replication across three Availability Zones (two active, one standby).
Amazon DocumentDB: Primary node and standby replica across two Availability Zones.
Compute tiers: The ingestion service and the frontend and search services use Auto Scaling groups distributed across Availability Zones, with the ALBs routing around failures.
Network isolation: Compute and data workloads reside in private subnets. Only the external ALB is publicly reachable, and outbound traffic flows through NAT gateways.
Results
Condé Nast conducted a benchmarking workshop in May 2026 to quantify impact. The following results reflect Condé Nast’s measurements from that workshop.
99.2 percent reduction in content discovery time, from 250 minutes to approximately 2 minutes per task (measured in the May 2026 benchmarking workshop described earlier).
Over 90 percent reduction in manual video review effort. Teams receive targeted clips with precise timestamps instead of scrubbing through individual videos (May 2026 benchmarking workshop described earlier).
Approximately $800,000 in estimated annual operational savings, based on productivity gains from alleviating repetitive manual search across the organization (estimate derived from the May 2026 benchmarking workshop described earlier).
Improved asset discoverability: Previously underutilized videos surfaced through semantic connections that metadata-only search could not identify.
Accelerated revenue capture: Faster discovery means faster response to advertiser requests and sales opportunities, shortening the path from inquiry to monetization.
“The AWS team, along with the GenAI and TwelveLabs attendees, helped us explain the business case clearly and concisely. I am optimistic about notching more progress towards our workflow goals in the near future.”
— Billy Keenly, Global Senior Director, Creative Optimization, Condé Nast
What we learned
Building this solution across an over 140,000 video library surfaced practical lessons that go beyond what architecture diagrams capture.
Early user research shaped the entire embedding and query design. The team interviewed editorial staff to understand how they describe content in their own words. Those conversations defined the right abstraction level for semantic search. Without that grounding, the system risked optimizing for queries no one actually types.
Separating ingestion from serving proved essential at this scale. Each plane could evolve, scale, and fail independently. The solution has been running in production for six months, and production search stayed available while the ingestion pipeline reprocessed backfill content or received updates.
Finding the right video segment length required deliberate experimentation. Short segments lost context. Long segments diluted the semantic signal. Iterative benchmarking against real editorial queries drove the team toward a duration that balanced precision and recall.
Asynchronous embedding generation was non-negotiable at this scale. Synchronous calls to the embedding model would have created bottlenecks across over 140,000 videos. Asynchronous invocation through AWS Step Functions allowed the pipeline to process the backlog without blocking, and it remains the pattern for ongoing ingestion.
Building in multi-AZ availability from day one avoided costly retrofits. For a solution that editorial teams rely on throughout the workday, even brief outages translate directly to lost productivity. It’s cheaper to build this in from the start than to add it later.
Conclusion
This post described how Condé Nast and AWS built a multimodal video intelligence solution using Amazon Bedrock, Amazon OpenSearch Service, and TwelveLabs Marengo. The solution replaces metadata-only search with intent-based semantic search across visual, audio, and transcript modalities. The result: a 99.2 percent reduction in discovery time and an estimated $800,000 in annual operational savings across an over 140,000 video library.
This pattern applies to organizations managing large video libraries, including broadcasters, streaming services, sports leagues, and enterprise media teams. The decoupled architecture adapts to different embedding models and content types as multimodal AI capabilities evolve.
To get started:
Learn more about Amazon Bedrock and explore the Amazon Bedrock User Guide for implementation guidance.
Explore Amazon OpenSearch Service and the k-NN vector search documentation for similarity search.
Read about TwelveLabs models on Amazon Bedrock for multimodal video understanding.
Review the AWS Step Functions documentation for orchestrated processing pipelines.
Contact the AWS Generative AI Innovation Center to discuss your use case or reach out to your AWS account team.
To get started with multimodal video discovery on AWS, visit Amazon Bedrock and Amazon OpenSearch Service. To learn how the AWS Generative AI Innovation Center can help your team build AI-powered solutions, visit the AWS Generative AI Innovation Center page.
Related posts
Unlocking video understanding with TwelveLabs Marengo on Amazon Bedrock
Multimodal embeddings at scale: AI data lake for media and entertainment workloads
Acknowledgements
We would like to thank the following contributors for their work on this solution: Shinan Zhang, Xiaoye Qian, Jessica Malca, Elvis Joseph, and Jason Janetzke.
About the authors