The same model can produce a vague, wrong answer or an excellent one depending on how you ask. Prompt engineering is the practice of designing inputs that reliably elicit the behaviour you need. It is less about magic phrases and more about clear communication plus empirical testing — skills that transfer across models.
Principle 1: be clear and specific#
Vague: "Summarise this report." Specific: "Summarise this report for a district health officer in five bullet points. Focus on water-borne disease risks and recommended actions. Use plain language and include any numbers mentioned."
State the task, the audience, the format, the length, the constraints and what to do when information is missing ("If the report does not mention X, say so").
Principle 2: give context#
Models do not know your situation. Provide the relevant background, definitions, policies or documents. When answers must come from specific sources, include them and instruct the model to rely on them — this is the idea behind retrieval-augmented generation.
Principle 3: use roles and system prompts#
A system prompt sets persistent behaviour: role, tone, scope, rules and safety constraints.
System: You are an assistant for a community health programme. Answer only using the provided
guidelines. If a question is outside them, reply "I don't have that information; please contact
the health post." Never give dosages for prescription medicines.Principle 4: show examples (few-shot prompting)#
Examples communicate format and edge cases better than descriptions. Choose diverse, representative examples, and include tricky ones.
Classify each message's urgency as HIGH, MEDIUM or LOW.
Message: "My child has had a high fever for three days and is now very drowsy."
Urgency: HIGH
Message: "When does the registration office open on Thursday?"
Urgency: LOW
Message: "Our tent roof is leaking and rain is expected tonight."
Urgency: MEDIUM
Message: "{new_message}"
Urgency:Be aware that models can be sensitive to example order and label balance.
Principle 5: ask for structured output#
For software integration, request JSON matching a schema — and validate it. Many APIs support structured outputs / JSON mode or function calling that constrains generation to a schema.
import json
from pydantic import BaseModel, ValidationError
class Extraction(BaseModel):
name: str | None
district: str | None
need: str
urgency: str
prompt = """Extract the fields from the message below. Respond ONLY with JSON:
{"name": string|null, "district": string|null, "need": string, "urgency": "HIGH"|"MEDIUM"|"LOW"}
Message: "This is Karim from Sylhet. We have had no clean water for two days."
"""
raw = '{"name": "Karim", "district": "Sylhet", "need": "clean water", "urgency": "HIGH"}' # model output
try:
data = Extraction(**json.loads(raw))
print(data)
except (json.JSONDecodeError, ValidationError) as e:
print("Invalid output, retry or route to human:", e)Principle 6: decompose complex tasks#
Break a hard task into steps — either within one prompt ("First list the key facts, then …") or across a chain of prompts (extract → verify → summarise). Asking the model to reason step by step before answering improves accuracy on many reasoning tasks (next lecture). For multi-step workflows, separate prompts are easier to test and debug.
Principle 7: reduce hallucination#
- Ask the model to answer only from provided context and to cite passages.
- Explicitly allow "I don't know".
- Ask for uncertainty or to list assumptions.
- Verify factual claims with retrieval or tools; do not rely on the model's memory for critical facts.
Principle 8: iterate empirically#
Treat prompts like code:
- Build a test set of realistic inputs (including edge cases and adversarial ones) with expected outputs or rubrics.
- Measure performance for each prompt version.
- Change one thing at a time; keep versions under source control.
- Re-test when the model changes — prompts are not guaranteed to transfer between models or versions.
def evaluate_prompt(template, cases, call_llm):
correct = 0
for case in cases:
output = call_llm(template.format(**case["inputs"]))
correct += case["check"](output) # e.g. exact label match or a rubric function
return correct / len(cases)Prompt injection#
Prompting vs fine-tuning#
Try prompting (and RAG) first — it is fast and cheap to iterate. Consider fine-tuning when you need consistent formats or styles across many requests, domain-specific behaviour that prompts cannot achieve, lower latency/cost (a small fine-tuned model replacing a large prompted one), or when prompts become unwieldy.