Languages🇸🇦AR🇩🇪DE🇫🇷FR
AI EngineeringJune 13, 2026·9 دقيقة قراءة·Miracle KaluMiracle Kalu

Building a Related-Articles Engine with Vector Embeddings: A Production-Ready Design

A glowing network of connected nodes representing vector embeddings and semantic similarity.

For years, the related-articles widget on our content site was a simple SQL query. It joined posts through shared tags, ordered the result by publication date, and called it a day. The implementation was cheap, predictable, and mostly irrelevant. Readers would finish an article about cache invalidation strategies and be offered the latest post tagged "DevOps" because someone had once used the word "deployment" in the introduction.

Tags describe editorial intent, not semantic meaning. They are great for navigation, but they are a blunt instrument for recommendation. Two articles can share every tag and still address completely different problems. Worse, two articles that solve the same problem in different domains may share no tags at all. As our archive grew past five hundred posts, the widget became a carousel of loosely related misses.

We needed a recommendation layer that understood what an article was actually about. Vector embeddings turned out to be the right tool, but only after we stopped treating them like a magic search box and started treating them like a production data pipeline.

What Embeddings Actually Buy You

An embedding is a dense vector that maps a piece of text into a high-dimensional semantic space. Texts with similar meanings end up close together, even if they use different words. That property gives you a recommendation signal that tags cannot: conceptual relatedness.

The model we chose sees "cache invalidation" and "keeping a CDN in sync with origin state" as neighbors. It sees "deployment pipelines" and "continuous delivery" as related without anyone having to tag them as such. The result is a related-articles widget that surfaces genuinely useful next reads.

Embeddings are not free. They cost money to generate, they take up storage, and they add latency to your query path. The production challenge is not generating the vectors. It is keeping the system fast, cheap, and maintainable while the archive grows.

Architecture Overview

The system has three independent paths: ingestion, query, and invalidation.

A common mistake is to put the embedding model directly on the request path. That couples reader traffic to model latency and provider availability. We generate embeddings at write time, store them in a vector database, and keep the query path as a simple nearest-neighbor lookup with a cache in front.

Choosing the Model and Provider

We started with OpenAI's text-embedding-3-small for three reasons: it is cheap, the context window covers most of our articles in a single pass, and the output dimension is configurable. For a recommendation widget, we found that 512 dimensions captured enough signal while keeping storage and query cost reasonable. Going to 1,536 dimensions improved quality only marginally and doubled our index size.

For teams with stricter data policies, self-hosted models like sentence-transformers/all-MiniLM-L6-v2 are a viable alternative. The trade-off is operational overhead. You now run GPU inference, manage model versioning, and monitor latency yourself. We chose the managed route because our volume did not justify the infrastructure cost, but the architecture works the same either way.

The important decision is not which provider you pick. It is isolating the provider behind an interface so you can swap it later without touching ingestion or query code.

The Ingestion Pipeline

Every time an article is published or updated, we run a background job that chunks the content, generates embeddings, and writes the vectors to the store. We do not embed the entire article as a single vector. Headlines, subheadings, and body paragraphs carry different semantic weight, and a single embedding of a three-thousand-word post tends to dilute the signal.

Our chunking strategy is simple but opinionated:

  • The title gets its own embedding.
  • Each H2 section gets its own embedding.
  • The body is split into paragraphs of up to 256 tokens, with a 32-token overlap.
  • Each chunk keeps a pointer back to the parent article, its locale, and its publish date.

This gives us multiple vectors per article, which improves recall. When a reader finishes a post, we query against all vectors for that article, aggregate the nearest neighbors by parent article, and rank the candidates by frequency and average distance.

interface ArticleChunk {
  id: string;
  articleId: string;
  locale: string;
  kind: "title" | "heading" | "paragraph";
  text: string;
  publishedAt: string;
}

interface EmbeddedChunk extends ArticleChunk {
  embedding: number[];
}

class EmbeddingPipeline {
  constructor(
    private embedder: EmbeddingProvider,
    private store: VectorStore,
  ) {}

  async ingest(article: Article): Promise<void> {
    const chunks = this.chunk(article);
    const embeddings = await this.embedder.embed(
      chunks.map((c) => c.text),
    );

    const withVectors: EmbeddedChunk[] = chunks.map((chunk, i) => ({
      ...chunk,
      embedding: embeddings[i],
    }));

    await this.store.upsert(`article:${article.id}`, withVectors);
  }

  private chunk(article: Article): ArticleChunk[] {
    const base = {
      articleId: article.id,
      locale: article.locale,
      publishedAt: article.publishedAt,
    };

    const chunks: ArticleChunk[] = [
      {
        id: `${article.id}:title`,
        ...base,
        kind: "title",
        text: article.title,
      },
      ...article.headings.map((h, i) => ({
        id: `${article.id}:h:${i}`,
        ...base,
        kind: "heading",
        text: h,
      })),
      ...splitParagraphs(article.body, 256, 32).map((p, i) => ({
        id: `${article.id}:p:${i}`,
        ...base,
        kind: "paragraph",
        text: p,
      })),
    ];

    return chunks;
  }
}

The splitParagraphs helper is intentionally naive. We split on sentence boundaries when possible, but we do not use sliding windows across sentences. For recommendation quality, preserving semantic boundaries matters more than maximizing token density.

Backfilling Without Blowing the Budget

The first time you turn this on, you will need to embed every article in your archive. If you have a thousand posts and each one makes several API calls, a naive backfill can get expensive fast. We used a rate-limited queue with cost tracking.

import PQueue from "p-queue";

async function backfill(
  articles: Article[],
  pipeline: EmbeddingPipeline,
  budgetCents: number,
): Promise<void> {
  const queue = new PQueue({ concurrency: 4 });
  const costPer1kTokens = 0.02; // USD
  let estimatedCost = 0;

  for (const article of articles) {
    const tokens = estimateTokens(article.body);
    estimatedCost += (tokens / 1000) * costPer1kTokens;

    if (estimatedCost * 100 > budgetCents) {
      console.warn("Backfill budget exhausted");
      break;
    }

    queue.add(() => pipeline.ingest(article));
  }

  await queue.onIdle();
}

Concurrency of four was the sweet spot for our embedding provider. Higher concurrency triggered rate limits without meaningfully improving throughput. We also ran the backfill during off-peak hours and logged every failure so we could retry specific articles without restarting the whole batch.

The query path is where caching and filtering become critical. A reader on the German version of an article should not see English recommendations. A reader on a post about frontend performance should not see backend database articles just because they share the word "query."

We query the vector store with all chunks from the current article, then aggregate candidate article IDs across the results. Each candidate gets a score based on how many of its chunks appeared in the top-k neighbors and how close those neighbors were. We also boost recent articles slightly so the widget does not always show five-year-old posts.

interface RelatedArticlesOptions {
  articleId: string;
  locale: string;
  category?: string;
  limit?: number;
}

class RelatedArticlesService {
  constructor(
    private store: VectorStore,
    private cache: Cache,
    private fallback: TagBasedFallback,
  ) {}

  async findRelated(opts: RelatedArticlesOptions): Promise<Article[]> {
    const cacheKey = `related:${opts.articleId}:${opts.locale}:${opts.category ?? "all"}`;
    const cached = await this.cache.get(cacheKey);
    if (cached) return cached;

    try {
      const chunks = await this.store.getArticleChunks(opts.articleId);
      if (chunks.length === 0) {
        return this.fallback.find(opts);
      }

      const candidates = await this.store.queryNeighbors(chunks, {
        excludeArticleId: opts.articleId,
        locale: opts.locale,
        category: opts.category,
        topK: 20,
      });

      const ranked = this.rank(candidates).slice(0, opts.limit ?? 5);
      await this.cache.set(cacheKey, ranked, { ttlSeconds: 3600 });
      return ranked;
    } catch (err) {
      console.error("Embedding query failed, falling back", err);
      return this.fallback.find(opts);
    }
  }

  private rank(candidates: Candidate[]): Article[] {
    const byArticle = new Map<string, Candidate[]>();
    for (const c of candidates) {
      const list = byArticle.get(c.articleId) ?? [];
      list.push(c);
      byArticle.set(c.articleId, list);
    }

    const scored = Array.from(byArticle.entries()).map(([articleId, hits]) => {
      const avgDistance =
        hits.reduce((sum, h) => sum + h.distance, 0) / hits.length;
      const recencyBoost = recencyScore(hits[0].publishedAt);
      return {
        articleId,
        score: hits.length * 0.6 + (1 - avgDistance) * 0.3 + recencyBoost * 0.1,
      };
    });

    return scored
      .sort((a, b) => b.score - a.score)
      .map((s) => ({ id: s.articleId }));
  }
}

The fallback to tag-based recommendations is not an admission of failure. It is a reliability guarantee. If the vector store is down, the embedding provider is rate-limiting you, or the article has no vectors yet, the widget still shows something relevant instead of an empty box or a 500 error.

Caching Strategy

The cache is the difference between a recommendation widget that adds tens of milliseconds and one that adds hundreds. We use Redis with a structured key and a stale-while-revalidate pattern for popular articles.

interface CacheEntry<T> {
  data: T;
  staleAt: number;
  expiresAt: number;
}

class RedisRelatedCache {
  constructor(private redis: RedisClient) {}

  async get<T>(key: string): Promise<T | null> {
    const raw = await this.redis.get(key);
    if (!raw) return null;

    const entry: CacheEntry<T> = JSON.parse(raw);
    if (Date.now() > entry.expiresAt) {
      await this.redis.del(key);
      return null;
    }

    return entry.data;
  }

  async set<T>(
    key: string,
    data: T,
    opts: { ttlSeconds: number; staleSeconds?: number },
  ): Promise<void> {
    const now = Date.now();
    const entry: CacheEntry<T> = {
      data,
      staleAt: now + (opts.staleSeconds ?? opts.ttlSeconds) * 1000,
      expiresAt: now + opts.ttlSeconds * 1000,
    };

    await this.redis.set(key, JSON.stringify(entry), "EX", opts.ttlSeconds);
  }

  async isStale(key: string): Promise<boolean> {
    const raw = await this.redis.get(key);
    if (!raw) return false;

    const entry: CacheEntry<unknown> = JSON.parse(raw);
    return Date.now() > entry.staleAt;
  }
}

Our cache keys include the article ID, locale, and optional category filter. We do not include user identity because the widget is the same for every reader of the same article. That keeps the cache hit rate high and the cardinality low.

TTL is one hour by default, with a stale window of five minutes. When a request finds a stale entry, we return it immediately and trigger a background refresh. This prevents cold articles from ever blocking a reader while keeping popular articles fresh.

Cache invalidation happens on publish and delete. The ingestion worker deletes the related-articles cache keys for the updated article and for any article that previously linked back to it. We do not try to be surgical. Invalidation is cheap; stale recommendations are expensive.

Filtering and Faceting

Raw vector similarity is not enough. A travel blog might have two articles about "Paris" that are semantically close but one is a budget guide and the other is a luxury hotel review. If your reader is on the budget guide, you probably do not want to send them to the luxury review.

We support two types of filters: hard filters and soft filters. Hard filters exclude candidates before the vector query returns, such as locale and publication status. Soft filters apply after ranking, such as category preference or reading time. Soft filters can be overridden if the semantic match is strong enough.

Vector stores vary widely in how well they support metadata filtering during ANN queries. Qdrant and Pinecone handle it well. PostgreSQL with pgvector works for smaller archives but struggles with combined vector and metadata queries at scale. We chose our store specifically because it could apply locale and category filters inside the ANN search, avoiding the need to fetch and filter large candidate sets afterward.

Monitoring and Cost Controls

Production systems that depend on a third-party API need guardrails. We track three metrics: embedding latency, query latency, and cost per thousand articles. Embedding latency is mostly a concern during backfills and bulk updates. Query latency is a reader-facing concern and is why the cache exists.

We also cap daily embedding spend. If a content migration or an import job suddenly enqueues tens of thousands of articles, we do not want a surprise invoice. The worker checks a daily budget counter in Redis before each batch and pauses when the cap is hit.

async function checkDailyBudget(
  redis: RedisClient,
  costCents: number,
  maxCents: number,
): Promise<boolean> {
  const key = `embed:budget:${new Date().toISOString().slice(0, 10)}`;
  const spent = await redis.incrby(key, Math.ceil(costCents));
  if (spent <= maxCents) {
    await redis.expire(key, 86_400);
  }
  return spent <= maxCents;
}

Finally, we log every query that falls back to tag-based recommendations. A sustained increase in fallback rate is usually the first sign of a vector store problem or a missing backfill.

Privacy and Data Retention

Sending article content to an embedding API has implications. Even though the content is already public, embedding providers may retain inputs for model improvement depending on their terms. We disable training use where the API allows it, and we audit the provider's data policy during contract review.

For internal or paywalled content, self-hosting is the safer default. The architecture does not change; only the provider interface does. That is why the abstraction matters.

Results and Caveats

After six weeks, click-through rate on the related-articles widget was up 34 percent. Time on site improved modestly. The biggest qualitative change was fewer complaints from editors that the widget was showing off-topic recommendations.

That said, embeddings are not always the right call. If your archive is small, a good full-text search index plus tags may give you better results with far less complexity. If your content is highly structured, entity-based recommendations can outperform vectors. Embeddings shine when the value is in the meaning of the text, not in explicit metadata.

They also require maintenance. Models get deprecated. Provider pricing changes. Vectors drift as your content strategy evolves. You are adding a new data pipeline, not just a widget.

Takeaways

  • Tags describe editorial intent; embeddings capture semantic meaning. Use the right signal for recommendations.
  • Generate embeddings at write time, not at request time. The query path should be a fast lookup with a cache in front.
  • Chunk content intelligently. Titles, headings, and paragraphs deserve separate vectors.
  • Always provide a fallback to tag-based or popularity-based recommendations. Reader-facing features cannot fail silently.
  • Cache aggressively with structured keys and stale-while-revalidate. Related articles do not need to be realtime.
  • Abstract the embedding provider and vector store behind interfaces. You will swap one of them eventually.
  • Monitor cost, latency, and fallback rate. These are the metrics that tell you if the system is healthy.
  • Treat embeddings as a data pipeline, not a plugin. It needs backfills, invalidation, budgeting, and retention policies.

مشاركة:

XLinkedIn
Miracle Kalu

كتبه

Miracle Kalu

Senior Full Stack Engineer

أعجبك ما قرأته؟

أنا متاح لأدوار الهندسة الأولى والاستشارات التقنية. لنتحدث.

تواصل →

نُشر 13 يونيو 2026 · 9 دقيقة قراءة

مواصلة القراءة