← All posts
How-ToAugust 28, 2026 · 12 min

Build Semantic Search for Legal Documents with Pinecone and OpenAI

You can build semantic search for legal documents without a brittle framework stack: use OpenAI to create embeddings, Pinecone to retrieve authorized passages, and the OpenAI Responses API to produce a source-grounded answer. The important part is not the demo query. It is the control layer around it—document permissions, tenant isolation, traceable citations, evaluation, and mandatory human review before anyone relies on an answer.

This blueprint uses current OpenAI and Pinecone Python SDK patterns, checked August 29, 2026. It is a research assistant, not legal advice and not a replacement for a qualified lawyer.

What you need

Components needed for secure legal semantic search
Components needed for secure legal semantic search
ComponentCurrent choiceRole
Python3.11+Ingestion and query service
OpenAI embeddingstext-embedding-3-small, 1,536 dimensionsConverts chunks and questions to vectors
OpenAI answer modelgpt-5.6-luna by defaultProduces a grounded answer from retrieved excerpts
PineconeServerless dense-vector indexStores vectors and metadata in isolated namespaces
PDF parserpypdf or your approved document pipelineExtracts text and page numbers
Application databaseYour existing Postgres or audit storeHolds users, permissions, document versions, and query logs

Install the small dependency set:

bash
python -m venv .venv
source .venv/bin/activate
pip install openai pinecone pypdf fastapi uvicorn

On Windows, activate with .venv\Scripts\activate. Store API keys in a secret manager or environment variables—never in source code or a browser bundle.

OpenAI currently lists text-embedding-3-small at $0.02 per million input tokens. Pinecone offers a Starter plan and paid tiers with different minimums and features. Treat all prices and rate limits as live configuration, not constants in your proposal: check the OpenAI model page and Pinecone pricing before estimating a client project.

1. Create a serverless Pinecone index

Lock the embedding dimension explicitly so a later model change cannot silently make stored and query vectors incompatible.

python
import os
from openai import OpenAI
from pinecone import Pinecone, ServerlessSpec

INDEX_NAME = "legal-docs"
DIMENSIONS = 1536

openai_client = OpenAI(api_key=os.environ["OPENAI_API_KEY"])
pc = Pinecone(api_key=os.environ["PINECONE_API_KEY"])

if not pc.has_index(INDEX_NAME):
    pc.create_index(
        name=INDEX_NAME,
        dimension=DIMENSIONS,
        metric="cosine",
        spec=ServerlessSpec(cloud="aws", region="us-east-1"),
    )

description = pc.describe_index(INDEX_NAME)
index = pc.Index(host=description.host)

Pinecone's current Python SDK package is named pinecone, not the retired pinecone-client package. The Pinecone Python SDK guide documents the Pinecone, ServerlessSpec, and host-targeting pattern used above.

2. Parse, chunk, and label every source

Legal retrieval fails when a chunk cannot be traced back to an authoritative page. Every vector should carry a stable document ID, source name, page, jurisdiction, version, and the text shown to the answer model.

python
from dataclasses import dataclass
import hashlib

@dataclass(frozen=True)
class LegalChunk:
    document_id: str
    source: str
    page: int
    jurisdiction: str
    text: str

    @property
    def vector_id(self) -> str:
        digest = hashlib.sha256(self.text.encode("utf-8")).hexdigest()[:16]
        return f"{self.document_id}#p{self.page}#{digest}"

Split on headings and paragraph boundaries first. A character- or token-based fallback can cap unusually long sections, but there is no universal “best” chunk size. Start with roughly 500–800 tokens and a small overlap, then measure retrieval on questions your legal team has labeled with the correct passages.

Do not ingest a folder merely because a service account can read it. Apply the source system's access-control list before extraction, preserve that authorization in your application database, and delete or re-index chunks when the source document or its permissions change.

3. Embed and upsert into an isolated namespace

Batch embeddings, keep the result order, and send traceable metadata with each vector:

python
def embed(texts: list[str]) -> list[list[float]]:
    response = openai_client.embeddings.create(
        model="text-embedding-3-small",
        dimensions=1536,
        input=texts,
    )
    return [item.embedding for item in sorted(response.data, key=lambda x: x.index)]

def upsert_chunks(namespace: str, chunks: list[LegalChunk]) -> None:
    vectors = []
    for chunk, values in zip(chunks, embed([c.text for c in chunks]), strict=True):
        vectors.append({
            "id": chunk.vector_id,
            "values": values,
            "metadata": {
                "document_id": chunk.document_id,
                "source": chunk.source,
                "page": chunk.page,
                "jurisdiction": chunk.jurisdiction,
                "text": chunk.text,
            },
        })
    index.upsert(namespace=namespace, vectors=vectors)

The namespace must come from the authenticated server-side account mapping. Never accept namespace or tenant_id directly from a public request body. Pinecone documents namespaces as an essential multitenancy mechanism, and its data-modeling guidance explains why physical namespace isolation also reduces the blast radius of filter bugs.

For larger imports, respect Pinecone's request-size limits and send bounded batches. Upserts are eventually consistent, so do not declare ingestion complete until your job verifies freshness. The current upsert guide covers batch limits and metadata behavior.

4. Retrieve with a permission boundary and metadata filter

Embed the question with the same model and dimensions, query only the authenticated tenant namespace, and add filters such as jurisdiction when the caller is authorized to use them.

python
def retrieve(namespace: str, question: str, jurisdiction: str, top_k: int = 5):
    query_vector = embed([question])[0]
    result = index.query(
        namespace=namespace,
        vector=query_vector,
        top_k=min(max(top_k, 1), 20),
        include_metadata=True,
        filter={"jurisdiction": {"$eq": jurisdiction}},
    )
    return [
        {
            "citation": f"S{position}",
            "score": match.score,
            "source": match.metadata["source"],
            "page": match.metadata["page"],
            "text": match.metadata["text"],
        }
        for position, match in enumerate(result.matches, start=1)
    ]

Do not treat similarity score as legal correctness. Establish a threshold from an evaluation set, show “not found” when retrieval is weak, and keep the original page available so a reviewer can inspect the full surrounding text.

5. Generate a fail-closed answer with verifiable citations

Send only the retrieved excerpts to the answer model. Require exact source labels and reject an answer that cites a label you did not provide.

python
import re

NOT_FOUND = "Not found in the authorized documents."

def answer_question(question: str, sources: list[dict]) -> str:
    if not sources:
        return NOT_FOUND

    context = "\n\n".join(
        f"[{s['citation']}] {s['source']}, page {s['page']}\n{s['text']}"
        for s in sources
    )
    response = openai_client.responses.create(
        model="gpt-5.6-luna",
        store=False,
        instructions=(
            "You are a legal research assistant, not a lawyer. Answer only from "
            "the supplied authorized excerpts. Cite every material claim with an "
            f"exact source label. If unsupported, reply exactly: {NOT_FOUND}"
        ),
        input=f"Question:\n{question}\n\nAuthorized excerpts:\n{context}",
    )
    answer = response.output_text.strip()
    if answer == NOT_FOUND:
        return answer

    allowed = {s["citation"] for s in sources}
    cited = set(re.findall(r"\[(S\d+)\]", answer))
    if not cited or not cited.issubset(allowed):
        raise ValueError("answer did not contain verifiable citations")
    return answer

store=False is a useful application-level data-control choice, but it does not replace a privacy review, contract terms, retention settings, regional requirements, or your organization's OpenAI data-control configuration. The OpenAI Responses API supports the response-generation flow; verify the model and data settings approved for your organization.

6. Put authentication and review in front of the API

Secure API flow for legal semantic search
Secure API flow for legal semantic search

Your HTTP endpoint should derive authorization from a verified session or service identity:

python
@app.post("/search")
def search(request: QueryRequest, user: User = Depends(require_user)):
    namespace = tenant_namespace_for(user)  # server-side lookup
    require_jurisdiction_access(user, request.jurisdiction)
    sources = retrieve(namespace, request.question, request.jurisdiction)
    answer = answer_question(request.question, sources)
    audit_log(user=user, question=request.question, sources=sources)
    return {"answer": answer, "sources": sources, "requires_human_review": True}

Add rate limiting, request-size caps, structured audit events, encryption in transit and at rest, source-document deletion workflows, and a review queue. Redact unnecessary personal data before embedding. For privileged or regulated material, involve counsel, security, privacy, and the client before any ingestion begins.

Pinecone plan features differ. Audit logs and service-account controls are not a substitute for your application's authorization and audit trail, and some platform controls are plan-dependent. Confirm the exact controls on Pinecone's live pricing and security documentation before making a compliance claim.

7. Evaluate before production

Create a versioned evaluation set with at least these fields:

  • question
  • authorized tenant and jurisdiction
  • expected source document and page
  • acceptable answer or “not found”
  • prohibited sources
  • reviewer decision

Measure retrieval recall, citation precision, unsupported-claim rate, cross-tenant leakage, “not found” behavior, latency, and cost. Include adversarial prompts that request another client's documents, instructions embedded inside a PDF, revoked documents, duplicate versions, scanned pages with poor OCR, and questions whose answer is absent.

Never promote a legal RAG release because five demo questions looked good. Gate the release on a representative, reviewer-approved dataset and re-run it whenever the embedding model, answer model, chunker, prompt, index settings, or source corpus changes.

Where this breaks

Cross-tenant leakage. A request-controlled namespace or metadata-only tenant filter can expose another client's documents. Fix it with server-derived namespaces, authorization tests, and deny-by-default source access.

Stale law or stale permissions. Vector search can faithfully retrieve a superseded case, revoked policy, or document the user no longer has permission to read. Version sources, record effective dates, and make permission changes trigger deletion or re-indexing.

Plausible unsupported answers. Retrieval reduces hallucination; it does not eliminate it. Enforce citation labels, validate them, show source pages, support “not found,” and require a lawyer or qualified reviewer.

Prompt injection in documents. A PDF can contain text telling the model to ignore its instructions or disclose data. Treat retrieved text as untrusted evidence, not instructions. Keep tools disabled for the answer step and test hostile documents.

Bad OCR and chunk boundaries. Tables, footnotes, exhibits, and multi-column scans can be extracted incorrectly. Preserve page images or source links, run OCR quality checks, and evaluate by document type.

Runaway cost and latency. Batch embeddings, cap top_k, monitor token and Pinecone read usage, and cache only when the cache key includes tenant, permissions, corpus version, and query settings.

FAQ

No. Legal research benefits from hybrid retrieval: semantic similarity finds paraphrases, while exact search protects citations, defined terms, docket numbers, and statutory language. Evaluate both and consider reranking a combined candidate set.

Can I use a different answer model?

Yes. Keep the model configurable and benchmark it on the same legal evaluation set. Model quality, price, limits, and availability change; a hard-coded marketing claim ages quickly.

No. Pinecone secures its service and offers plan-dependent controls, but your application must authenticate users, authorize documents, select the correct namespace, record reviewer-visible audit events, and enforce retention and deletion policies.

How do I update a document?

Use deterministic document and chunk IDs. When a document changes, delete the old document's chunks in its namespace, upsert the new version, verify freshness, and record the corpus version used by each answer.

No. It is a software architecture pattern for authorized legal research. A qualified professional must validate the sources, answer, and compliance design for the actual jurisdiction and use case.

Next steps

Start with one authorized corpus and 30–50 reviewer-labeled questions. Build the deny-by-default permission layer first, then ingest, retrieve, cite, and evaluate. Only after the leakage tests and citation gate pass should you add a polished interface.

For more implementation patterns, explore AI automations you can sell and get the free AI automation guide.

Get the full toolkit

Grab the free guide with the node-by-node build for all 10 automations.

No spam. Unsubscribe anytime. Just the good stuff.