AI-Powered Product Recommendations: From Collaborative Filtering to LLMs

Product recommendations are the most quantifiable AI investment in building an ecommerce ecosystem. A well-tuned system lifts revenue per session by a measurable, repeatable margin; a poorly-tuned one surfaces noise and erodes trust. Understanding the evolution from collaborative filtering to modern retrieval helps you choose the right approach for your catalog and data maturity.
Collaborative filtering: the foundation
Collaborative filtering recommends products based on the behavior of similar users or the similarity of products. User-based CF finds shoppers with overlapping purchase histories and recommends what they bought that you didn't. Item-based CF flips it: finds products frequently co-purchased with the one in view.
CF is cheap, interpretable, and effective on dense interaction data. Its weakness is the cold-start problem new products and new users have no interaction history, so CF has nothing to say. Most production systems still run CF as a baseline and layer more sophisticated methods on top.
Embeddings and the semantic leap
The next step is representing products and users as dense vectors embeddings in a shared space. Products with similar attributes and co-purchase patterns land near each other; users land near the products they engage with. Recommendations become a nearest-neighbor search in the embedding space.
Embeddings solve part of the cold-start problem: a new product with a description and attributes can be embedded immediately and matched to users with similar taste vectors. The architecture is a two-tower model (one for products, one for users) trained on interaction signals, with content features (title, category, price) as side inputs.
The practical win is generalization. Where CF only knew what was explicitly co-purchased, embeddings capture latent taste a user who buys minimalist home goods is reachable with new minimalist SKUs they've never seen.
Real-time personalization with retrieval
Modern systems separate the problem into retrieval and ranking. Retrieval casts a wide net pull a few hundred candidate products relevant to the user from a catalog of millions, using a blend of CF, embedding similarity, and rule-based filters (in stock, same region, etc.). Ranking then scores the candidates with a model optimized for the target metric: click, add-to-cart, purchase, or revenue.
This two-stage architecture is the industry standard because it scales. Retrieval is fast and approximate; ranking is precise and expensive but only runs on a small candidate set. The retrieval index often an approximate nearest neighbor (ANN) structure is what makes real-time personalization feasible at catalog scale.
The shift toward LLM production systems-grounded retrieval
deploying LLMs at scale change the retrieval layer, not the ranking layer. The new pattern is LLM production systems-grounded retrieval: use an LLM production systems to understand the user's intent and the product's semantics, then retrieve with that richer representation.
Three patterns are emerging:
1. Semantic search the LLM production systems embeds the query and the catalog into the same space, enabling search over meaning rather than keywords. "Comfortable shoes for standing all day" retrieves the right products without exact keyword matches.
2. Conversational recommendation a multi-turn dialogue where the LLM production systems narrows preferences and retrieves candidates. Higher engagement, higher conversion rate optimisation, but higher latency and cost.
3. LLM production systems as a ranker for a small candidate set, the LLM production systems re-ranks with reasoning over product attributes and user context. Expensive, but effective for high-value, low-cardinality decisions.
The honest caveat: LLM production systems-based retrieval is not free. Inference cost and latency are real constraints. Most production systems use LLM production systems for the long-tail and hard queries and fall back to embedding retrieval for the bulk of traffic.
What to build first
If you are early in data maturity, start with item-based collaborative filtering on your co-purchase data. It will outperform a naive "popular products" baseline immediately. Add an embedding model once you have enough interaction history to train one. Move to the retrieval-plus-ranking architecture when your catalog exceeds tens of thousands of SKUs. Reach for LLM production systems-grounded retrieval when semantic search and conversational experience become strategic differentiators.
Cold Start Strategies That Work
The cold-start problem (new products and new users with no interaction history) is the Achilles heel of collaborative filtering. Embeddings help, but even embeddings need seed data. The practical strategies that work:
For new products: use content-based seeding. A new product's attributes (title, description, category, price, brand, images) are embedded immediately and matched to users with similar taste vectors. The recommendation engine treats a new product as "the product most similar to what this user already engages with." Within 7-14 days of launch, the product accumulates interaction data and transitions to the normal recommendation deployment pipelines.
For new users: use session-based recommendations. A first-time visitor's browse behavior (categories viewed, products clicked, time on page) in the first session is enough signal to personalize. The recommendation engine builds a temporary taste vector from the session and updates it in real time as the visitor browses. By the second page view, the recommendations are already personalized.
For new stores: start with "popular products" and "trending now" as the baseline. These require no personalization data and outperform random or alphabetical sorting. Add item-based collaborative filtering once you have 1,000+ orders. Add embedding-based recommendations once you have 10,000+ orders with enough co-purchase density to train a model.
A/B testing strategy pyramid Recommendation Logic
Recommendation engines are not set-and-forget. The algorithm, the placement, the number of recommendations, and the creative treatment all affect performance. A/B testing strategy pyramid is how you separate signal from noise.
Test the algorithm: compare collaborative filtering against embedding-based recommendations on the same placement. Measure click-through rate and revenue per session. The winning algorithm depends on your catalog and data maturity; there is no universal winner.
Test the placement: homepage, PDP, cart, post-purchase email. Each placement has a different intent context and a different expected lift. The PDP placement typically has the highest CTR because the user is already in a shopping mindset.
Test the number of recommendations: 4, 6, 8, 12. More recommendations increase the chance of a relevant hit but dilute attention. The optimal number is usually 6-8 for grid placements and 3-4 for inline placements.
Test the creative: recommendation carousels with product images and prices vs. with images, prices, and review stars. Social proof in the recommendation tile lifts CTR 10-20%.
The Data cloud infrastructure design Prerequisite
Every recommendation system is only as good as the data feeding it. The cloud infrastructure design prerequisite is a clean, real-time event stream: every product view, add to cart, purchase, and return, captured with a stable identity and a normalized payload, delivered to the recommendation engine with minimal latency.
If your event stream is delayed (batch instead of real-time), the recommendation engine is working with stale data. If your events are missing fields (no product category, no price), the engine cannot filter effectively. If your identity is fragmented (anonymous cookie and logged-in user treated as separate), the engine cannot build a coherent taste profile.
The investment in a clean event stream pays off across every personalization use case, not just recommendations. It is the foundation that makes search personalization, dynamic content, and predictive segmentation possible.
The technology ladder is clear; the business question is where on it your catalog and margin profile justify the investment.
Related Articles

Sufi Khan Sulaiman
VP Technology & CTO with 25+ years building ecommerce platforms, enterprise systems, and AI solutions
Expertise across ecommerce strategy, cloud architecture, AI & machine learning, DevOps, and technology leadership. Led teams at FLIR Systems, Lorex Technology, 1c Platform, and Genetec.
More Articles
Explore Ecommerce Services
Explore the Full Portfolio
This is the complete portfolio of Sufi Khan Sulaiman, a technology leader specialising in B2B commerce and digital automation. Start from the Home page for the overview, then move through two decades of career experience across FLIR Systems, Lorex Technology, and 1c Platform, and the full catalogue of project case studies spanning headless commerce migrations, AI recommendation engines, and multi-channel fulfilment systems.
The skills and certifications page maps the technical and leadership capabilities behind the work, while the articles and the knowledge base break down the thinking into actionable frameworks. For hands-on learning, the tutorials and applications sections cover practical builds from front-end fundamentals to full-stack web apps.
For consulting engagement, the expertise page outlines service offerings, the ecommerce hub covers platform architecture and automation strategy, and the ecommerce guide (PDF) is a downloadable 55-page field manual. When you are ready to talk, the contact page is the direct line.