SASERÁ Academy
Generative AI · concept guide

RAG vs fine-tuning: the decision in one page

Both make a foundation model more useful for your business. One gives it the right documents at the right moment; the other changes the model itself. Here is how to choose.

AWS AI Practitioner exam prep cover
RAGRetrieval Augmented Generation
  • Model weights stay the same
  • Answers use current documents at question time
  • Can cite the sources behind each answer
  • Update knowledge by updating documents
  • On AWS: Amazon Bedrock Knowledge Bases
Fine-tuningTrain the model on your examples
  • Model weights change
  • Teaches style, format, tone or a narrow task
  • Needs labeled prompt-completion examples
  • Must retrain when the desired behavior changes
  • On AWS: Bedrock model customization or SageMaker AI

Decision guide

Facts change weekly or dailyRAG
Answers must show their sourcesRAG
A fixed output format or brand voiceFine-tuning
Lots of unlabeled domain text, new vocabularyContinued pre-training
A smaller, cheaper model for one narrow taskDistillation
A quick win with no infrastructurePrompt engineering first

From cheapest to most expensive

  1. 1

    Prompt engineering

    Clear instructions, examples and output format. No extra infrastructure.

  2. 2

    RAG

    Adds embeddings, a vector store and retrieval, but no training.

  3. 3

    Fine-tuning

    Adds labeled data preparation and training jobs, repeated when behavior must change.

  4. 4

    Pre-training from scratch

    Massive data, compute and expertise. Rarely the right business choice.

Exam tip: when a question says the information "changes frequently" or must be "up to date", the answer is almost never fine-tuning.

4 practice questions: RAG or fine-tuning?

Taken from the AIF-C01 course. Click an option; every option is explained.

0 of 4 answered · 0 correctEvery option is explained after you answer
Fundamentals of GenAIQuestion 1 of 4

A pharmaceutical company's assistant answers questions about internal standard operating procedures that are updated weekly. The company wants answers to reflect the latest version without retraining any model. Which approach BEST meets this requirement?

  • Temperature cannot give the model knowledge of documents it has never seen.
  • Correct. RAG grounds each answer in the latest documents, so weekly updates flow through without retraining.
  • Weekly fine-tuning adds cost and delay, and it is what the company wants to avoid.
  • Pre-training is far too expensive and slow for weekly document changes.
Why it matters: Retrieval-augmented generation looks up current documents at query time and passes them to the model. Updating the knowledge source updates answers without changing model weights.
Fundamentals of GenAIQuestion 2 of 4

A consulting firm wants its GenAI assistant to answer questions using its private project reports and to show which report each answer came from. Which approach BEST meets both needs?

  • Fine-tuned models do not reliably reproduce exact sources, so citations may be invented.
  • Correct. Retrieved chunks carry source metadata, which the application can show as citations.
  • Bigger models know more, so this is tempting. The firm's reports are private and were never in any pre-training data, and pre-trained knowledge cannot point to which report an answer came from.
  • The base model has never seen the firm's private reports and cannot cite them.
Why it matters: RAG retrieves relevant passages from private sources and passes them to the model, and the retrieved sources can be returned as citations for verification.
Applications of Foundation ModelsQuestion 3 of 4

A two-person startup wants a foundation model to return customer feedback summaries in a fixed three-part format. It has almost no budget and only a handful of example summaries. Which approach should the startup try FIRST?

  • A few examples are too little data for effective fine-tuning, and training adds cost.
  • It teaches domain vocabulary from unlabeled text but not a three-part output format, and it costs far more than a startup with almost no budget can spend.
  • RAG adds facts from a data source; it needs infrastructure and is not the cheapest way to control format. Placing the examples directly in the prompt is simpler.
  • Correct. Including a few examples in the prompt (few-shot, in-context learning) often teaches format with no training cost and can be tested immediately.
Why it matters: In-context learning changes behavior through the prompt only, making it the lowest-cost customization approach to try first.
Fundamentals of GenAIQuestion 4 of 4

A company needs a model to answer accurately in a narrow domain and is weighing fine-tuning against RAG. From a cost perspective, which statement is MOST accurate?

  • Changing knowledge sounds like it needs retraining, so this is tempting. RAG only re-indexes documents; it is fine-tuning that must be repeated to update the model's knowledge.
  • Owning a custom model sounds cheaper to run, so this is tempting. Custom models still incur inference charges, and fine-tuning adds training and hosting or storage costs.
  • Embedding feels like a one-time step, so this is tempting. RAG keeps costing money for the vector store, query embeddings, and the extra retrieved tokens in every prompt.
  • Correct. Each approach shifts cost to a different place, so the choice depends on the need.
Why it matters: Fine-tuning adds training, storage, and often dedicated hosting costs, and must be repeated as data changes. RAG adds retrieval and extra input tokens but avoids retraining.

Customization questions are 28% of AIF-C01

Drill them in six timed exams where every option is explained.

See the AIF-C01 course →

FAQ

What is the difference between RAG and fine-tuning?

RAG retrieves relevant documents at question time and adds them to the prompt, without changing the model. Fine-tuning changes the model's weights by training it on labeled examples.

Which is cheaper?

RAG is usually cheaper to start and to keep current, because updating documents is cheaper than retraining. Fine-tuning adds training jobs and must be repeated when the desired behavior changes.

When is fine-tuning the better choice?

When you need a consistent style, format or tone, or a specialized task the model does poorly even with good prompts, and the knowledge itself does not change often.

Can I use both?

Yes. Many production systems fine-tune for style or format and use RAG for current facts.

More guides: AIF-C01 study guide · Bedrock vs SageMaker AI · AI Practitioner vs Cloud Practitioner · Security+ 30-day plan · all guides