跳到主要内容
AI News HubLIVE
来源内容 · 翻译待补全6 分钟阅读

待翻译:Uplifting conversion across the acquisition funnel with personalization using contextual bandits on AWS

文章摘要

AI 服务暂时不可用,以下为来源摘要,待恢复后补全翻译:Generative AI makes it cheap to produce personalized content at scale, but which variation do you show each customer? Amazon Payments used a multi-objective contextual bandit on Amazon SageMaker AI to personalize an acquisition funnel, achieving a high single-digit conversion lift for one audience, and learning why content, not the model, was the constraint.

来源AWS Machine Learning Blog作者: Chidi Prince John
待翻译:Uplifting conversion across the acquisition funnel with personalization using contextual bandits on AWS
报告错误

纠错通道尚未开通,可先复制下方文章信息留存。

查看更正说明
直接读正文

AI 服务暂时不可用,以下为来源正文,待恢复后补全翻译。

Generative AI has made it possible to produce large amounts of personalized content quickly and at low cost. In our previous post, we showed how generative AI on Amazon Bedrock can produce personalized content at scale while staying within brand guidelines and guardrails. The new challenge is now one of selection. Among all of those options, which one do you show each customer, and how long does it take to learn the answer? This post tackles the selection challenge that follows the proliferation. In this post, we share how Amazon Payments applied AI-based personalization to a product acquisition funnel, using a multi-objective contextual multi-armed bandit (MAB) on Amazon SageMaker AI. In a seven-week online A/B test we currently see a high single-digit percentage relative lift in final-funnel conversion for one customer population, while another saw no improvement over the existing experience. The problem turned out to be the content, not the model. We cover the intuition behind bandits, our extension to optimize an entire conversion funnel, and the AWS architecture behind the solution. We also share a code repository that you can use to test this approach on synthetic data and understand the method hands-on using Amazon SageMaker AI. Growing role of multi-armed bandits in the generative AI era A multi-armed bandit (MAB) is a reinforcement learning method built for settings with many options and limited traffic. It learns which option performs best while continuing to serve customers. It treats each content variation as an “arm,” tries each against live traffic, and steadily shifts impressions toward the arms that perform, while holding a fraction back to keep testing the rest. This is the fundamental trade-off between exploitation (serve the current best arm) and exploration (try less-certain arms to gather evidence). Because a bandit never stops doing both, it keeps improving as new variations are added, and it never has to wait for a test to conclude. A/B/n testing still has its place, but as generative AI accelerates the number of variations to learn from, we expect bandits to play a growing role. There are several bandit selection strategies in the literature: epsilon-greedy, Upper Confidence Bound (UCB), Thompson sampling, among others. For a deeper introduction to these methods and their deployment on AWS, see our earlier post Dynamic A/B testing for machine learning models with Amazon SageMaker MLOps Projects. In Amazon Payments, we chose UCB, a strategy that selects the arm with the highest estimated reward plus an uncertainty bonus, naturally balancing exploitation and exploration. Its deterministic selection rule gives an auditable, reproducible decision for every impression, which makes sure every serving decision can be explained and reproduced if needed. However, a standard bandit learns one best arm for the entire audience. Personalization requires conditioning on who the visitor is, and that is what contextual bandits provide. Adding the context using contextual bandits Instead of asking “which content is the best overall?,” a contextual bandit conditions its decision on signals about the visit. Segmented bandits run a separate instance per hand-defined group, but the groups are arbitrary and each needs its own traffic. Contextual bandits condition directly on a feature vector (a numeric representation of attributes), so a pattern learned in one context transfers to every similar visit without separate per-group traffic. Our production system represents each customer as a context vector of behavioral signals (payment behavior, transaction mix, and similar features) in place of a fixed segment. The entity ID (an opaque key such as entity_id) is used only to route the learned recommendation back to the right visitor. It is never a model input. LinUCB: Learning at the feature level For a contextual approach, we selected Linear UCB (LinUCB), introduced by Li et al. (2010). While newer bandit algorithms exist, LinUCB remains a battle-tested method that is computationally efficient, auditable (deterministic arm selection), and naturally handles a large arm space with limited warm-start data. Its key assumption is that the expected reward for an arm is a linear function of the context vector, which lets it generalize to visitors it has not seen before. Each arm keeps two running tallies updated on every impression: b, the reward ledger: which visitor signals led to conversions (b += reward · x). A, the experience ledger: which visitors the arm has seen (A += x·xᵀ, starting from identity). Dividing reward by experience gives the estimate (θ = A⁻¹·b). The experience ledger also shrinks the exploration bonus as evidence grows. The arm’s score is: score(a, x) = θᵀ·x + α · √(xᵀ · A⁻¹ · x) \___/ \____________/ estimate uncertainty bonus The bonus is context-dependent. It’s large for a kind of visitor the arm has rarely seen, and small for one it has seen often. A plain bandit explores at a single global rate, while LinUCB adjusts how much it explores for every visitor it scores. Optimizing an entire funnel The customer journey in our use case comprises three steps: application start, submission, and approval. The objective is therefore a multi-stage outcome, not a single metric. These stages do not move in unison. Content optimized for starts tends to attract a broad audience, yet approval depends on whether the offer genuinely suits the applicant. Optimizing one stage in isolation can degrade another. We refer to this as the seesaw problem. Conversely, optimizing solely for approvals starves the model of signal, because approvals are rare and delayed. Our approach optimizes the entire funnel simultaneously by running one LinUCB model per stage (start, submit, approve) and combining their UCB scores through a linear combination: selected_arm = argmax_a [ w_start · UCB_start(a, x) + w_submit · UCB_submit(a, x) + w_approve · UCB_approve(a, x) ] The stage weights can be assigned based on business priorities or learned by a separate calibration step. We used approximately equal weights. In practice, one could weight approvals more heavily once the model is warmed or use a Pareto frontier if the trade-off is genuinely contested. Here’s the core of the model in Python. Each LinUCBDisjoint maintains independent parameters per arm, and MultiObjectiveLinUCB composes three of them: import numpy as np class LinUCBDisjoint: """LinUCB with disjoint (independent) parameters per arm. Each arm maintains its own reward and experience matrices.""" def init(self, n_arms, n_features, alpha=1.0): self.n_arms = n_arms self.n_features = n_features self.alpha = alpha # exploration-exploitation trade-off parameter # A: experience matrix (d x d); b: reward vector (d) - one per arm self.A = [np.eye(n_features) for _ in range(n_arms)] self.b = [np.zeros(n_features) for _ in range(n_arms)] def update(self, arm, context, reward): """Exact, incremental update - no retraining needed.""" self.A[arm] += np.outer(context, context) self.b[arm] += reward * context class MultiObjectiveLinUCB: """Combines one LinUCB per funnel stage via weighted sum.""" def init__(self, n_arms, n_features, alpha=1.0, weights=(1/3, 1/3, 1/3)): # stage weights (start, submit, approve) self.n_arms = n_arms self.alpha = alpha self.weights = weights # [start, submit, approve] self.app_start_bandit = LinUCBDisjoint(n_arms, n_features, alpha) self.app_submit_bandit = LinUCBDisjoint(n_arms, n_features, alpha) self.app_approved_bandit = LinUCBDisjoint(n_arms, n_features, alpha) def select_arm(self, context): combined = [] for a in range(self.n_arms): score = 0.0 for w, bandit in zip( self.weights, [self.app_start_bandit, self.app_submit_bandit, self.app_approved_bandit], ): A_inv = np.linalg.inv(bandit.A[a]) theta = A_inv @ bandit.b[a] ucb = theta @ context + self.alpha * np.sqrt( context @ A_inv @ context) score += w * ucb combined.append(score) return int(np.argmax(combined)) def update(self, arm, context, r_start, r_submit, r_approve): self.app_start_bandit.update(arm, context, r_start) self.app_submit_bandit.update(arm, context, r_submit) self.app_approved_bandit.update(arm, context, r_approve) We built a quick start Jupyter notebook to test this approach on Amazon SageMaker AI. The repository also includes a command-line demo and unit tests. Tuning the alpha. The α parameter controls the exploration-exploitation balance. Higher values encourage exploration of under-tested arms, and lower values favor exploitation. A value of α = 1.0 is a reasonable default (Li et al., 2010). In practice, start higher when the arm space is large and history is limited, then reduce α as evidence accumulates. A typical range is 0.1 to 2.0. Values that are too high waste traffic on weak arms, and values that are too low risk locking onto a suboptimal arm. Because α is a single scalar, a grid search is straightforward. Handling delayed feedback. In our setting, approval decisions lag by days, so the reward for the final funnel stage isn’t observed until well after the impression. We address this with an attribution window. Starts and submissions update the model immediately, while approval outcomes are held back until the subsequent batch cycle, avoiding downward bias from pending applications. This aligns naturally with our weekly processing cadence. Publishing safe content at scale A bandit is only as good as the pool of arms. Building a rich arm pool responsibly is the next problem to solve at scale while maintaining content oversight. The key idea is to compose many variations from a small set of reviewed building blocks. Our content was assembled from two kinds of parts: industry-themed images and benefit-focused taglines. The arm space is the Cartesian product of these parts, so each arm is one (image, tagline) pairing and a modest number of building blocks produces a large pool of distinct page variations. Figure 1: Building blocks are vetted individually, then combined into the full arm pool That structure also makes the approach scalable without sacrificing oversight, through a few layers of control. Vetting the parts, not the combinations. Some content elements are fixed by stakeholder requirements. Rather than reviewing every possible combination, which would be impractical as the space grows, we vet each individual building block up front. Because the number of building blocks remains small and reviewable while the combinations grow rapidly, vetting the parts once provides assurance over the entire combinatorial space. Design systems for consistency. As described in our earlier post, anchoring content to a design system (approved colors, layouts, components, and patterns) keeps variations visually consistent and on-brand by construction, not by chance. Expanding the pool with generative AI. In our initial deployment, building blocks were curated with detailed human oversight, even when production was assisted by generative AI tools. We see clear potential to increase the parts pool (text, imagery, layouts) substantially, which would widen the bandit’s selection space. As this pool grows, the bandit and the generative pipeline form a virtuous cycle. Generative AI widens the candidate set, and the bandit identifies which combinations yield the strongest outcomes per visitor. Deploying on AWS Amazon SageMaker AI is an AI model development service that supports the full model lifecycle. It’s the primary AWS service behind our end-to-end pipeline, providing model training, batch processing, version control, and monitoring in a unified tool. We adopted a batch architecture for two reasons: the offline nature of the selection problem (feedback accumulates over days), and the observation that visitor selection behavior does not shift rapidly enough to require real-time model updates. A scheduled SageMaker AI Processing job reads the prior period’s feedback, [truncated for AI cost control]

展开要点与分析

文章情报

投资人中级

要点

  • AI 服务暂时不可用,系统已先保留来源内容与降级元数据。
  • Generative AI makes it cheap to produce personalized content at scale, but which variation do you show each customer? Amazon Payments used a multi-objective contextual bandit on A…

要点与分析由自动化流程生成,可能有误,请结合原始来源核实。