✨ Generative AI & LLMs · Lecture 15 of 30

Prompt Engineering: Getting Reliable Results from LLMs

Practical, evidence-based techniques for prompting LLMs — clear instructions, context, examples, output formats, decomposition, and systematic evaluation — plus prompt injection risks.

The same model can produce a vague, wrong answer or an excellent one depending on how you ask. Prompt engineering is the practice of designing inputs that reliably elicit the behaviour you need. It is less about magic phrases and more about clear communication plus empirical testing — skills that transfer across models.

Principle 1: be clear and specific#

Vague: "Summarise this report." Specific: "Summarise this report for a district health officer in five bullet points. Focus on water-borne disease risks and recommended actions. Use plain language and include any numbers mentioned."

State the task, the audience, the format, the length, the constraints and what to do when information is missing ("If the report does not mention X, say so").

Principle 2: give context#

Models do not know your situation. Provide the relevant background, definitions, policies or documents. When answers must come from specific sources, include them and instruct the model to rely on them — this is the idea behind retrieval-augmented generation.

Principle 3: use roles and system prompts#

A system prompt sets persistent behaviour: role, tone, scope, rules and safety constraints.

text
System: You are an assistant for a community health programme. Answer only using the provided
guidelines. If a question is outside them, reply "I don't have that information; please contact
the health post." Never give dosages for prescription medicines.

Principle 4: show examples (few-shot prompting)#

Examples communicate format and edge cases better than descriptions. Choose diverse, representative examples, and include tricky ones.

text
Classify each message's urgency as HIGH, MEDIUM or LOW.

Message: "My child has had a high fever for three days and is now very drowsy."
Urgency: HIGH

Message: "When does the registration office open on Thursday?"
Urgency: LOW

Message: "Our tent roof is leaking and rain is expected tonight."
Urgency: MEDIUM

Message: "{new_message}"
Urgency:

Be aware that models can be sensitive to example order and label balance.

Principle 5: ask for structured output#

For software integration, request JSON matching a schema — and validate it. Many APIs support structured outputs / JSON mode or function calling that constrains generation to a schema.

python
import json
from pydantic import BaseModel, ValidationError

class Extraction(BaseModel):
    name: str | None
    district: str | None
    need: str
    urgency: str

prompt = """Extract the fields from the message below. Respond ONLY with JSON:
{"name": string|null, "district": string|null, "need": string, "urgency": "HIGH"|"MEDIUM"|"LOW"}

Message: "This is Karim from Sylhet. We have had no clean water for two days."
"""
raw = '{"name": "Karim", "district": "Sylhet", "need": "clean water", "urgency": "HIGH"}'  # model output
try:
    data = Extraction(**json.loads(raw))
    print(data)
except (json.JSONDecodeError, ValidationError) as e:
    print("Invalid output, retry or route to human:", e)

Principle 6: decompose complex tasks#

Break a hard task into steps — either within one prompt ("First list the key facts, then …") or across a chain of prompts (extract → verify → summarise). Asking the model to reason step by step before answering improves accuracy on many reasoning tasks (next lecture). For multi-step workflows, separate prompts are easier to test and debug.

Principle 7: reduce hallucination#

  • Ask the model to answer only from provided context and to cite passages.
  • Explicitly allow "I don't know".
  • Ask for uncertainty or to list assumptions.
  • Verify factual claims with retrieval or tools; do not rely on the model's memory for critical facts.

Principle 8: iterate empirically#

Treat prompts like code:

  1. Build a test set of realistic inputs (including edge cases and adversarial ones) with expected outputs or rubrics.
  2. Measure performance for each prompt version.
  3. Change one thing at a time; keep versions under source control.
  4. Re-test when the model changes — prompts are not guaranteed to transfer between models or versions.
python
def evaluate_prompt(template, cases, call_llm):
    correct = 0
    for case in cases:
        output = call_llm(template.format(**case["inputs"]))
        correct += case["check"](output)          # e.g. exact label match or a rubric function
    return correct / len(cases)

Prompt injection#

Prompting vs fine-tuning#

Try prompting (and RAG) first — it is fast and cheap to iterate. Consider fine-tuning when you need consistent formats or styles across many requests, domain-specific behaviour that prompts cannot achieve, lower latency/cost (a small fine-tuned model replacing a large prompted one), or when prompts become unwieldy.

JA
Written by

Janin A Apurba

B.Sc. in CSE, AUST · Advanced ICT Officer, CNRS-UNHCR. Teaching AI, ML and Deep Learning to the next generation of engineers and researchers.

Keep learning

Related lectures

✨ Generative AI & LLMs

In-Context Learning: How LLMs Learn from Prompts

Large language models can perform new tasks from a few examples in the prompt without weight updates. We examine what in-context learning is, what influences it, theories of how it works, and its practical limits.

Intermediate⏱ 5 min#207
✨ Generative AI & LLMs

LoRA and Parameter-Efficient Fine-Tuning (PEFT)

Full fine-tuning of billion-parameter models is expensive. PEFT methods train a tiny fraction of parameters. We derive LoRA's low-rank updates, QLoRA's 4-bit training, compare adapters and prompt tuning, and give practical recipes.

Advanced⏱ 5 min#210
✨ Generative AI & LLMs

Large Language Models: What They Are and How They Are Built

A map of large language models — the transformer backbone, the training pipeline from pretraining to alignment, what capabilities emerge, how they are served and used, and their fundamental limitations.

Beginner⏱ 5 min#199