HardКейс14 min

Рекомендательные системы

Collaborative filtering, content-based, hybrid подходы и проектирование рекомендательной системы

Задача

Спроектировать рекомендательную систему для e-commerce платформы с 10M пользователей и 1M товаров. Система должна генерировать персонализированные рекомендации в реальном времени.

Требования

Требование Значение
Latency < 100ms для online рекомендаций
Throughput 10,000 RPS
Freshness Учитывать последние действия (< 1 час)
Coverage Рекомендовать > 80% каталога
Diversity Не повторять одни и те же товары

Подходы к рекомендациям

Сравнение подходов

Подход Идея Плюсы Минусы
Collaborative Filtering Похожие пользователи любят похожее Не нужен контент Cold start
Content-Based Рекомендовать похожее на то, что нравилось Нет cold start для items Filter bubble
Hybrid Комбинация подходов Лучшее качество Сложность
Knowledge-Based Правила и ограничения Прозрачность Не масштабируется
Deep Learning Neural networks State-of-the-art Требует данных и GPU

Collaborative Filtering

User-Based CF

Найти пользователей с похожими вкусами и рекомендовать то, что нравится им.

User A: [Item1: 5, Item2: 3, Item3: ?, Item4: 4]
User B: [Item1: 4, Item2: 3, Item3: 5, Item4: ?]
User C: [Item1: 2, Item2: 1, Item3: 2, Item4: 1]

A и B похожи → рекомендовать Item3 для A (B поставил 5)

Item-Based CF

Найти товары, которые часто покупают вместе.

Users who bought Item1 also bought: Item3 (80%), Item5 (65%), Item7 (40%)

Item-Based Collaborative Filtering

<?php

declare(strict_types=1);

namespace App\Recommendation;

final class ItemBasedCF
{
    /**
     * Calculate item-item similarity using cosine similarity.
     *
     * @param array<string, array<string, float>> $userItemMatrix
     *        [user_id => [item_id => rating]]
     * @return array<string, array<string, float>>
     *        [item_id => [similar_item_id => similarity_score]]
     */
    public function calculateItemSimilarity(array $userItemMatrix): array
    {
        // Transpose: item → users who rated it
        $itemUsers = [];

        foreach ($userItemMatrix as $userId => $items) {
            foreach ($items as $itemId => $rating) {
                $itemUsers[$itemId][$userId] = $rating;
            }
        }

        $items = array_keys($itemUsers);
        $similarity = [];

        for ($i = 0; $i < count($items); $i++) {
            for ($j = $i + 1; $j < count($items); $j++) {
                $score = $this->cosineSim(
                    $itemUsers[$items[$i]],
                    $itemUsers[$items[$j]],
                );

                if ($score > 0.1) { // Only store meaningful similarities
                    $similarity[$items[$i]][$items[$j]] = $score;
                    $similarity[$items[$j]][$items[$i]] = $score;
                }
            }
        }

        return $similarity;
    }

    /**
     * Get recommendations for a user.
     *
     * @param array<string, float> $userRatings [item_id => rating]
     * @param array<string, array<string, float>> $itemSimilarity
     * @param int $topN Number of recommendations
     * @return array<array{item_id: string, score: float}>
     */
    public function recommend(
        array $userRatings,
        array $itemSimilarity,
        int $topN = 10,
    ): array {
        $scores = [];
        $simSums = [];

        foreach ($userRatings as $ratedItem => $rating) {
            $similar = $itemSimilarity[$ratedItem] ?? [];

            foreach ($similar as $candidateItem => $similarity) {
                // Skip items user already rated
                if (isset($userRatings[$candidateItem])) {
                    continue;
                }

                $scores[$candidateItem] = ($scores[$candidateItem] ?? 0) + $similarity * $rating;
                $simSums[$candidateItem] = ($simSums[$candidateItem] ?? 0) + $similarity;
            }
        }

        // Normalize scores
        $recommendations = [];

        foreach ($scores as $itemId => $score) {
            if ($simSums[$itemId] > 0) {
                $recommendations[] = [
                    'item_id' => $itemId,
                    'score' => round($score / $simSums[$itemId], 4),
                ];
            }
        }

        // Sort by score descending
        usort($recommendations, static fn($a, $b) => $b['score'] <=> $a['score']);

        return array_slice($recommendations, 0, $topN);
    }

    private function cosineSim(array $a, array $b): float
    {
        $commonUsers = array_keys(array_intersect_key($a, $b));

        if (empty($commonUsers)) {
            return 0.0;
        }

        $dot = 0.0;
        $normA = 0.0;
        $normB = 0.0;

        foreach ($commonUsers as $user) {
            $dot += $a[$user] * $b[$user];
            $normA += $a[$user] ** 2;
            $normB += $b[$user] ** 2;
        }

        $denom = sqrt($normA) * sqrt($normB);

        return $denom > 0 ? $dot / $denom : 0.0;
    }
}
## Content-Based Filtering

Рекомендовать товары, похожие на те, что пользователь уже оценил.

<?php

declare(strict_types=1);

namespace App\Recommendation;

final readonly class ContentBasedRecommender
{
    public function __construct(
        private ProductRepository $products,
    ) {}

    /**
     * Recommend similar products based on attributes.
     *
     * @param array<string> $likedProductIds Products user liked
     * @param int $limit Max recommendations
     * @return array<array{product_id: string, score: float}>
     */
    public function recommend(array $likedProductIds, int $limit = 10): array
    {
        // Build user profile from liked items
        $userProfile = $this->buildUserProfile($likedProductIds);

        // Score all candidate products
        $candidates = $this->products->findAll();
        $scores = [];

        foreach ($candidates as $product) {
            if (in_array($product['id'], $likedProductIds, true)) {
                continue; // Skip already liked
            }

            $productFeatures = $this->extractFeatures($product);
            $score = $this->similarity($userProfile, $productFeatures);

            if ($score > 0.1) {
                $scores[] = [
                    'product_id' => $product['id'],
                    'score' => round($score, 4),
                ];
            }
        }

        usort($scores, static fn($a, $b) => $b['score'] <=> $a['score']);

        return array_slice($scores, 0, $limit);
    }

    /**
     * Build user preference profile from liked items.
     */
    private function buildUserProfile(array $productIds): array
    {
        $aggregated = [];
        $count = 0;

        foreach ($productIds as $id) {
            $product = $this->products->findById($id);

            if ($product === null) {
                continue;
            }

            $features = $this->extractFeatures($product);

            foreach ($features as $key => $value) {
                $aggregated[$key] = ($aggregated[$key] ?? 0) + $value;
            }

            $count++;
        }

        // Average
        if ($count > 0) {
            foreach ($aggregated as &$value) {
                $value /= $count;
            }
        }

        return $aggregated;
    }

    /**
     * Extract numerical features from a product.
     */
    private function extractFeatures(array $product): array
    {
        return [
            'price_normalized' => min($product['price'] / 10000, 1.0),
            'category_' . ($product['category'] ?? 'other') => 1.0,
            'brand_' . ($product['brand'] ?? 'other') => 1.0,
            'rating' => ($product['rating'] ?? 0) / 5.0,
            'popularity' => min(($product['sales_count'] ?? 0) / 1000, 1.0),
        ];
    }

    private function similarity(array $a, array $b): float
    {
        $allKeys = array_unique(array_merge(array_keys($a), array_keys($b)));
        $dot = 0.0;
        $normA = 0.0;
        $normB = 0.0;

        foreach ($allKeys as $key) {
            $va = $a[$key] ?? 0.0;
            $vb = $b[$key] ?? 0.0;
            $dot += $va * $vb;
            $normA += $va ** 2;
            $normB += $vb ** 2;
        }

        $denom = sqrt($normA) * sqrt($normB);

        return $denom > 0 ? $dot / $denom : 0.0;
    }
}
## Архитектура Production-системы
[API Request] → [Recommendation Service]
                       ↓
               ┌───────────────┐
               │   Candidate   │ ← Pre-computed candidates (batch)
               │   Generation  │ ← User history (online)
               └───────┬───────┘
                       ↓
               ┌───────────────┐
               │   Ranking     │ ← ML model (online prediction)
               │   Model       │
               └───────┬───────┘
                       ↓
               ┌───────────────┐
               │   Post-       │ ← Dedup, diversity, business rules
               │   Processing  │ ← Remove out-of-stock, seen items
               └───────┬───────┘
                       ↓
               [Final Recommendations]

Two-stage подход

Этап Задача Latency Кандидаты
Candidate Generation Выбрать ~1000 кандидатов < 20ms 1M → 1000
Ranking Ранжировать кандидатов < 50ms 1000 → 50
Post-processing Фильтры, разнообразие < 10ms 50 → 10

Cold Start Problem

Тип Проблема Решения
New User Нет истории действий Популярные товары, демографика, onboarding quiz
New Item Нет рейтингов/покупок Content-based, editorial curation, boost in ranking
New System Нет данных вообще Правила, тренды, экспертные рекомендации

Метрики рекомендательных систем

Метрика Формула Что измеряет
Precision@K Relevant in top-K / K Релевантность
Recall@K Relevant in top-K / Total relevant Полнота
NDCG Normalized DCG Качество ранжирования
CTR Clicks / Impressions Вовлечённость
Coverage Recommended items / All items Охват каталога
Diversity 1 - avg similarity between recs Разнообразие
Novelty Avg popularity rank of recs Неожиданность

Итоги

Концепция Суть
Collaborative Filtering Похожие пользователи / похожие товары
Content-Based Рекомендовать похожее по атрибутам
Hybrid Комбинация для лучшего качества
Two-stage Candidate generation → ranking
Cold Start Решается через fallback стратегии
A/B Testing Единственный способ проверить quality

Принцип: Простая модель, развёрнутая в production с хорошим A/B тестированием, всегда лучше совершенной модели, которая живёт в notebook.