30-SECOND SUMMARY
What to take away
- Create 10–20 Korean prompts with known evidence.
- Fix runtime, prompt, temperature, context, and quantization.
- Score fluency and factual accuracy separately.
Compare Korean output
Separate fluency from facts.
- 01SET
Fixed Korean cases
- 02REPEAT
Same conditions
- 03SCORE
Facts, language, format
- 04SELECT
Check critical errors
Break down the work
Summarization, extraction, honorific register, spacing, and structured output are different capabilities. Weight the set according to real use.
The working rule for “Break down the work” is: Create 10–20 Korean prompts with known evidence. Keep the model, quantization, context length, and concurrency fixed so that a device or model comparison has a clear cause.
For verification, save the model and runtime versions, source input, relevant settings, and observed output together. Repeat the step while changing only one factor, and record unexpected results and untested limits as carefully as successes before applying the guidance to private or production data.
Define expected evidence
List required facts and disqualify unsupported claims. Use public or anonymized documents.
The working rule for “Define expected evidence” is: Fix runtime, prompt, temperature, context, and quantization. Keep the model, quantization, context length, and concurrency fixed so that a device or model comparison has a clear cause.
Preserve the before-and-after state and the time of the check so that another run can reproduce the result. Include at least one failure condition—such as empty input, constrained resources, or a restart—to reveal the boundary of the step rather than documenting only the happy path.
Repeat under fixed conditions
Run each case multiple times and record the full model name, version, runtime, and date.
The working rule for “Repeat under fixed conditions” is: Score fluency and factual accuracy separately. Keep the model, quantization, context length, and concurrency fixed so that a device or model comparison has a clear cause.
Define completion with an observable result instead of a general impression. Repeat the same input, and if the output changes, isolate whether the model, runtime settings, or source data changed before moving to the next stage.
Score separate dimensions
Measure factual match, instruction following, natural Korean, format validity, consistency, and speed.
The working rule for “Score separate dimensions” is: Create 10–20 Korean prompts with known evidence. Keep the model, quantization, context length, and concurrency fixed so that a device or model comparison has a clear cause.
Treat the table as a recording framework, not as a universal answer. Fill it with measurements from the intended device and workload, and mark unavailable values as unknown instead of replacing them with zero or an estimate that could distort the comparison.
| Dimension | Check |
|---|---|
| Facts | Names and numbers match |
| Summary | No key omission or invention |
| Language | Consistent terms and register |
| Format | Requested table or JSON is valid |
Frequently asked questions
Does a strong English benchmark guarantee Korean quality?
No. Test the Korean workflow directly. For a practical check, follow the “Break down the work” section, change one condition at a time, and record the result.
Is translation testing enough?
No. Include native Korean extraction, summary, and register. For a practical check, follow the “Define expected evidence” section, change one condition at a time, and record the result.
How many repetitions?
Three per prompt is a useful starting point. Score fluency and factual accuracy separately. For a practical check, follow the “Repeat under fixed conditions” section, change one condition at a time, and record the result.
Primary sources
Check the original documentation for version-specific details.
Ollama Chat API LM Studio model guide