Your question is LLM Fine-Tuning vs Prompting Tradeoffs. Take a moment with it on the right.
Talk me through your thinking if you like. When you're confident, submit your answer and I'll grade it like a real screen (7/10 or better passes).
You need an LLM for a text classification task and are deciding between two paths: fine-tuning a smaller open-source model or prompting a larger hosted model. You want to compare them in a practical way, not just by raw accuracy.
What are the trade-offs of fine-tuning a smaller open-source LLM versus prompting a larger proprietary model for classification tasks?