Paul SerbanSoftware Engineer
  • PORTFOLIO
  • BLOG

#LLM Architecture

Posts

  • Architecting Production LLM Systems: How AI Gateways, Orchestration, RAG, and Observability Fit TogetherA practical reference architecture connecting OpenRouter, LangChain/LlamaIndex, vector search, structured outputs, and Langfuse into one coherent AI engineering stack

    How AI gateways, LangChain/LlamaIndex orchestration, vector search, structured outputs, and Langfuse observability combine into a production LLM architecture.

    • #AI Engineering
    • #LLM Architecture
    • #Retrieval Augmented Generation
    • #AI Gateway
    • #LangChain
  • CAP theorem in AI systems: the consistency-availability trade-off your LLM pipeline is already makingHow distributed systems theory maps onto the core tensions in AI engineering

    Explore how CAP theorem principles from distributed systems map to LLM pipeline design, revealing the hidden trade-offs in consistency, availability, and reliability.

    • #LLM Architecture
    • #System Design
    • #AI Engineering
    • #CAP Theorem
    • #Reliability
  • Command vs. query prompts: a CQRS-inspired framework for structuring LLM interactionsTreating action-oriented and retrieval-oriented prompts as fundamentally different concerns leads to cleaner, more predictable AI behaviour

    Apply CQRS thinking to prompt engineering - learn how splitting command and query prompts improves LLM reliability, tracing, and eval coverage.

    • #AI Engineering
    • #LLM Architecture
    • #Prompt Engineering
    • #CQRS
    • #System Design
View all posts
  • LinkedInLinkedIn
  • GitHubGitHub
  • HackerRankHackerRank
  • LeetCodeLeetCode
  • EmailEmail
  • Portfolio
  • My Projects
  • Coursework
  • Blog
  • Posts
  • Snippets
  • Book Notes
  • Cookie Settings
  • Cookie Policy
2026 © Paul Serban. All rights reserved.www.paulserban.eu | Sitemap