RAG vs fine-tuning: the decision in one page
Both make a foundation model more useful for your business. One gives it the right documents at the right moment; the other changes the model itself. Here is how to choose.

- Model weights stay the same
- Answers use current documents at question time
- Can cite the sources behind each answer
- Update knowledge by updating documents
- On AWS: Amazon Bedrock Knowledge Bases
- Model weights change
- Teaches style, format, tone or a narrow task
- Needs labeled prompt-completion examples
- Must retrain when the desired behavior changes
- On AWS: Bedrock model customization or SageMaker AI
Decision guide
From cheapest to most expensive
- 1
Prompt engineering
Clear instructions, examples and output format. No extra infrastructure.
- 2
RAG
Adds embeddings, a vector store and retrieval, but no training.
- 3
Fine-tuning
Adds labeled data preparation and training jobs, repeated when behavior must change.
- 4
Pre-training from scratch
Massive data, compute and expertise. Rarely the right business choice.
4 practice questions: RAG or fine-tuning?
Taken from the AIF-C01 course. Click an option; every option is explained.
A pharmaceutical company's assistant answers questions about internal standard operating procedures that are updated weekly. The company wants answers to reflect the latest version without retraining any model. Which approach BEST meets this requirement?
- Temperature cannot give the model knowledge of documents it has never seen.
- Correct. RAG grounds each answer in the latest documents, so weekly updates flow through without retraining.
- Weekly fine-tuning adds cost and delay, and it is what the company wants to avoid.
- Pre-training is far too expensive and slow for weekly document changes.
A consulting firm wants its GenAI assistant to answer questions using its private project reports and to show which report each answer came from. Which approach BEST meets both needs?
- Fine-tuned models do not reliably reproduce exact sources, so citations may be invented.
- Correct. Retrieved chunks carry source metadata, which the application can show as citations.
- Bigger models know more, so this is tempting. The firm's reports are private and were never in any pre-training data, and pre-trained knowledge cannot point to which report an answer came from.
- The base model has never seen the firm's private reports and cannot cite them.
A two-person startup wants a foundation model to return customer feedback summaries in a fixed three-part format. It has almost no budget and only a handful of example summaries. Which approach should the startup try FIRST?
- A few examples are too little data for effective fine-tuning, and training adds cost.
- It teaches domain vocabulary from unlabeled text but not a three-part output format, and it costs far more than a startup with almost no budget can spend.
- RAG adds facts from a data source; it needs infrastructure and is not the cheapest way to control format. Placing the examples directly in the prompt is simpler.
- Correct. Including a few examples in the prompt (few-shot, in-context learning) often teaches format with no training cost and can be tested immediately.
A company needs a model to answer accurately in a narrow domain and is weighing fine-tuning against RAG. From a cost perspective, which statement is MOST accurate?
- Changing knowledge sounds like it needs retraining, so this is tempting. RAG only re-indexes documents; it is fine-tuning that must be repeated to update the model's knowledge.
- Owning a custom model sounds cheaper to run, so this is tempting. Custom models still incur inference charges, and fine-tuning adds training and hosting or storage costs.
- Embedding feels like a one-time step, so this is tempting. RAG keeps costing money for the vector store, query embeddings, and the extra retrieved tokens in every prompt.
- Correct. Each approach shifts cost to a different place, so the choice depends on the need.
Customization questions are 28% of AIF-C01
Drill them in six timed exams where every option is explained.
FAQ
What is the difference between RAG and fine-tuning?
RAG retrieves relevant documents at question time and adds them to the prompt, without changing the model. Fine-tuning changes the model's weights by training it on labeled examples.
Which is cheaper?
RAG is usually cheaper to start and to keep current, because updating documents is cheaper than retraining. Fine-tuning adds training jobs and must be repeated when the desired behavior changes.
When is fine-tuning the better choice?
When you need a consistent style, format or tone, or a specialized task the model does poorly even with good prompts, and the knowledge itself does not change often.
Can I use both?
Yes. Many production systems fine-tune for style or format and use RAG for current facts.
More guides: AIF-C01 study guide · Bedrock vs SageMaker AI · AI Practitioner vs Cloud Practitioner · Security+ 30-day plan · all guides