Is Your AI Failing? How to Use DSPy in Python to Fix Prompts

Learn how to use dspy in python to replace manual prompt engineering with declarative programs and automatic optimizers in 2026.

Is Your AI Failing? How to Use DSPy in Python to Fix Prompts
Source (Personal archive/maiastudios.com.br)

If you've spent hours tweaking adjectives in a system prompt only to watch your application break after a model update, you know traditional prompt engineering doesn't scale in production. Treating language models as black boxes that require handcrafted text tweaks introduces fragility and blocks serious software versioning. Understanding how to use dspy in python resolves this bottleneck by turning LLM-driven development into a declarative programming discipline, where code defines task structure and optimization algorithms automatically compile the best instructions.

In this practical tutorial, we'll ditch fragile string formatting hacks and build a robust pipeline using DSPy in Python 3.14.8. You'll see how to define typed signatures, chain reasoning modules, optimize prompts with training datasets, and deploy reproducible artificial intelligence programs into production.

What is DSPy and why did manual prompt engineering fail?

DSPy (Declarative Self-improving Python), developed by Stanford University's NLP group, fundamentally changes how we interact with language models. Instead of writing massive prompt strings filled with manual few-shot examples inside your code, developers express desired behavior using declarative signatures and reusable modules.

Manual prompt engineering failed because it is extremely sensitive to minor variations. When OpenAI or Anthropic updates a model, a perfectly tuned prompt can suddenly output invalid JSON or hallucinate parameters. Furthermore, as pipeline complexity grows—integrating vector search (RAG), schema validation, and external tools—managing dependencies across dozens of prompt strings becomes unsustainable.

DSPy introduces the concept of prompt compilation. Just as a C compiler converts high-level source code into machine instructions optimized for a specific processor, DSPy's optimizer takes your Python signatures, runs evaluations against a dataset, and compiles the ideal prompt for your chosen model (whether it's a proprietary API model or a local instance running on vLLM).

How to use dspy in python to structure and optimize AI calls?

Abstract vector illustration on a dark background showing the compilation flow of a Python code signature transforming input data through multiple optimization blocks to the final output.
Source (Personal archive/maiastudios.com.br)

To understand how to use dspy in python in your project, the first step is setting up the environment and connecting the framework to your chosen model provider. Under the hood, DSPy uses LiteLLM to standardize communication across hundreds of inference providers.

Start by installing the official library in your virtual environment:

pip install dspy

After installation, configure the global language model using the dspy.LM class. You can supply commercial API models or local endpoints compatible with the OpenAI API:

import dspy

# Global language model configuration
lm = dspy.LM('openai/gpt-4o-mini', api_key='sk-sua-chave-aqui')
dspy.configure(lm=lm)

# Quick connectivity test
resposta = lm('Explique o que é tipagem estática em uma frase.')
print(resposta)

The major advantage of this abstraction is portability. If you decide to swap your cloud provider for a local LLM running on your own server in 2026, you only need to change the initialization string in dspy.LM. Everything else in your codebase, including validation and extraction logic, remains completely untouched.

How do Signatures and Modules work in DSPy?

The two core pillars of DSPy are Signatures and Modules. A Signature defines the input and output contract of an AI task without specifying how the prompt should be visually formatted.

You can define a Signature using simplified inline syntax or by subclassing dspy.Signature to add explanatory docstrings and explicit types via Pydantic:

from pydantic import BaseModel, Field

# Defining the input and output contract
class ClassificarBug(dspy.Signature):
    """Analisa um relatório de erro e extrai o componente afetado e a gravidade."""

    relatorio_erro: str = dspy.InputField(desc="Descrição técnica do bug enviada pelo usuário")
    componente: str = dspy.OutputField(desc="Módulo do sistema afetado (ex: banco_dados, auth, interface)")
    nivel_gravidade: str = dspy.OutputField(desc="Classificação: ALTA, MEDIA ou BAIXA")

With the Signature defined, you instantiate a Module. DSPy provides several pre-built modules that implement well-known reasoning patterns:

  • dspy.Predict: Executes a direct call based on the signature.
  • dspy.ChainOfThought: Forces the model to generate an intermediate reasoning step (rationale) before answering.
  • dspy.ReAct: Enables the model to use external tools in decision loops.

Here is how to use the ChainOfThought module with our custom signature:

# Instanciando o módulo de raciocínio encadeado
classificador = dspy.ChainOfThought(ClassificarBug)

# Executando a inferência
resultado = classificador(
    relatorio_erro="Falha de Timeout ao tentar conectar no PostgreSQL 18.6 durante o pico de acesso."
)

print(f"Raciocínio: {resultado.rationale}")
print(f"Componente: {resultado.componente}")
print(f"Gravidade: {resultado.nivel_gravidade}")

DSPy automatically generates the prompt template with clear field boundary instructions, parses the output, and returns a typed object. If you want to inspect the exact prompt sent to the model, simply call dspy.inspect_history(n=1).

How do you compile programs with MIPROv2 and BootstrapFewShot optimizers?

The real magic of DSPy happens when you stop making isolated calls and start compiling your programs. Compilation is the process where an optimization algorithm (called a Teleprompter or Optimizer) evaluates your program against a small training dataset and automatically selects the best instructions and few-shot examples.

To compile a program, you need three components:

  1. The DSPy program (composed of one or more modules).
  2. A validation metric (a Python function that returns True/False or a numerical score).
  3. A set of training examples (dspy.Example).

Let's build a complete optimizer pipeline using MIPROv2 (Multistage Instruction Proposal and Optimization), one of the most advanced optimizers available in the framework in 2026:

# 1. Defining the training dataset
treino = [
    dspy.Example(relatorio_erro="NullPointer ao autenticar token JWT expirado", componente="auth", nivel_gravidade="MEDIA").with_inputs('relatorio_erro'),
    dspy.Example(relatorio_erro="Vazamento de memória no worker Celery consumindo 16GB RAM", componente="infraestrutura", nivel_gravidade="ALTA").with_inputs('relatorio_erro'),
    dspy.Example(relatorio_erro="Botão de salvar perfil com desalinhamento de 2px no CSS", componente="interface", nivel_gravidade="BAIXA").with_inputs('relatorio_erro'),
]

# 2. Defining the evaluation metric
def metrica_avaliacao(example, pred, trace=None):
    componente_correto = example.componente.lower() == pred.componente.lower()
    gravidade_correta = example.nivel_gravidade.upper() == pred.nivel_gravidade.upper()
    return componente_correto and gravidade_correta

# 3. Configuring the MIPROv2 optimizer
otimizador = dspy.MIPROv2(
    metric=metrica_avaliacao,
    auto="light" # Configuration for fast optimization with few iterations
)

# 4. Compiling the program
programa_compilado = otimizador.compile(
    student=dspy.ChainOfThought(ClassificarBug),
    trainset=treino,
    max_bootstrapped_demos=2,
    max_labeled_demos=2
)

# 5. Saving the optimized program to disk
programa_compilado.save("classificador_bugs_otimizado.json")

During compile execution, MIPROv2 tests different system instruction variations and picks the best examples from successful run histories. The final output is a lightweight JSON file containing the exact prompts that scored highest on your metric.

To load and run the program in production, you don't need to re-run the optimizer. The loading process is instantaneous:

programa_producao = dspy.ChainOfThought(ClassificarBug)
programa_producao.load("classificador_bugs_otimizado.json")

resultado_final = programa_producao(relatorio_erro="Erro 500 ao emitir nota fiscal via webhook")
print(resultado_final.componente)

How does DSPy compare to traditional orchestration frameworks?

A common software architecture question is understanding where DSPy fits compared to established alternatives like LangChain and Pydantic AI. The key difference isn't in connectivity features, but in design philosophy.

While traditional frameworks focus on abstractions for chaining calls and integrating tools, DSPy focuses on the mathematical and algorithmic optimization of instructions sent to the model.

Comparison Criterion DSPy LangChain Pydantic AI
Primary Approach Declarative programming and compilation Chaining and integration ecosystem Strict data validation in Python
Prompt Creation Automatic via search algorithms Manual via string interpolation Manual with Jinja / docstrings
Maintainability High: changes require re-compilation Medium: prompts scattered across code High: tightly integrated with Pydantic
Adapting to New Models Automatic upon re-compiling against dataset Manual: rewrite prompts for each LLM Manual: adjust prompts in code
Learning Curve Medium (requires a mindset shift) Low initial, high complexity at scale Low for developers already using Pydantic

Note that DSPy doesn't exclude validation libraries. In fact, you can use Pydantic schemas directly in DSPy Signatures to enforce strict type constraints on returned data.

How do you monitor and deploy DSPy programs in production?

Shallow depth of field photograph of an artificial intelligence development server with GPUs illuminated by cyan and violet LEDs in a dark studio.
Source (Personal archive/maiastudios.com.br)

Deploying a DSPy application to a production environment requires similar care to any critical microservice. Because the compiled program is saved as a static artifact (the JSON file with selected instructions and demonstrations), you should treat this file with the same rigor reserved for binaries or model weights.

Check the compiled file into version control (Git) or store it in an artifact repository during your CI/CD pipeline. Then, expose inference using a high-performance asynchronous framework like FastAPI or Litestar running under Python 3.14.8:

from fastapi import FastAPI, HTTPException
from pydantic import BaseModel
import dspy

app = FastAPI(title="Serviço de Triagem de Bugs")

# Initialization and loading on service startup
lm = dspy.LM('openai/gpt-4o-mini')
dspy.configure(lm=lm)

classificacao_app = dspy.ChainOfThought(ClassificarBug)
classificacao_app.load("classificador_bugs_otimizado.json")

class SolicitacaoBug(BaseModel):
    relatorio: str

class RespostaBug(BaseModel):
    componente: str
    gravidade: str
    raciocinio: str

@app.post("/triagem", response_model=RespostaBug)
async def triar_bug(payload: SolicitacaoBug):
    try:
        resposta = classificacao_app(relatorio_erro=payload.relatorio)
        return RespostaBug(
            componente=resposta.componente,
            gravidade=resposta.nivel_gravidade,
            raciocinio=resposta.rationale
        )
    except Exception as e:
        raise HTTPException(status_code=500, detail=str(e))

For monitoring cost and latency, set up tracking hooks with open observability tools like OpenTelemetry. Since DSPy calls pass through LiteLLM, capturing per-request token usage, validation error counts, and time-to-first-token (TTFT) metrics is straightforward.

Another best practice is setting up a periodic re-compilation pipeline. Whenever the support team collects new production data from resolved tickets, add those examples to your validation dataset and schedule an async job to re-compile the program. This way, your system continuously improves as usage grows—without any developer having to write a single line of prompt text manually.

Conclusion

Handcrafted prompt engineering is on its way out in enterprise software development. As applications require greater reliability and structured responses, relying on manual text tweaks becomes an unacceptable architectural risk. Mastering how to use dspy in python allows you to treat LLM calls like traditional software components: built with well-defined interfaces, testable with objective metrics, and automatically compilable.

By adopting Signatures, Modules, and Optimizers in your projects, you build pipelines that evolve alongside language models, ensuring predictability, traceability, and high performance in production.

Enjoyed it? Share

More in Innovation & Trends