Reviewed 6 September 2026

Part I: Domains

LLMs, agents and retrieval

Library Description
openai, anthropic, google-genai First-party API clients.
LiteLLM Unified interface and proxy across many providers.
Instructor, Outlines Structured output: schema-constrained generation and validation.
PydanticAI Type-safe agent framework from the Pydantic team. V2, released June 2026, is a breaking change from V1: it moves configuration onto a composable “capability” primitive and splits fast-moving pieces into a separate Harness package beside a slimmer core.
LangGraph Graph-based orchestration with explicit state, checkpointing and human-in-the-loop steps.
OpenAI Agents SDK Agents, handoffs, guardrails and tracing. Provider-agnostic in practice, still on 0.x.
Claude Agent SDK Anthropic’s agent framework, including subagent spawning.
LlamaIndex Indexing, retrieval and query pipelines. Workflows is a separate package, llama-index-workflows, on 2.x since August 2025.
mcp, FastMCP Model Context Protocol SDK and a higher-level server framework for building MCP tools.
pgvector Vector column type and index for PostgreSQL.
Qdrant, Chroma, LanceDB Vector stores; Rust-backed with payload filtering, embedded/client-server, and Lance-columnar respectively.
sentence-transformers Embedding and reranking model inference.
rank_bm25 Lexical BM25 scoring, used for hybrid retrieval.
Langfuse, Logfire Tracing and evaluation for LLM applications.
Ragas, DeepEval Evaluation metrics for retrieval and generation quality.

Example designs

Document question answering

  1. Ingestion
  2. structural chunking
  3. sentence-transformers embeddings
  4. pgvector in existing PostgreSQL
  5. hybrid retrieval (vector + rank_bm25)
  6. reranker
  7. generation with citations

Storing vectors in the operational database removes a second system and keeps chunks transactionally consistent with their source documents. Chunking follows the document’s own headings and table boundaries rather than a fixed character count, which keeps tables intact. Hybrid retrieval covers cases where the query contains exact identifiers that embeddings handle poorly, such as part numbers and error codes. A held-out set of question and expected-source pairs is run as pytest cases with Ragas metrics, so retrieval changes are measured rather than assessed by inspection.

Typed extraction service

  1. FastAPI endpoint
  2. PydanticAI agent with an output model
  3. LiteLLM provider routing
  4. validated object

The output schema is a Pydantic model, and validation failures trigger a bounded retry with the error fed back to the model. Requests carry a schema version so downstream consumers can handle changes. Prompt and model identifiers are logged with each response, because a silent provider-side model update is otherwise indistinguishable from a regression in your own code.