FICTIONAL LLM EVALUATION SAMPLE

Transparent scoring. Reproducible findings.

A deterministic Python harness that ranks two fictional AI answers against visible requirements, evidence records, and prohibited claims. It makes no model calls and produces the same structured report from the same input.

TYPED FINDINGS

Connect every deduction to observable evidence.

中文摘要:该虚构样例用明确权重评估指令遵循、事实支持与完整性,并输出可复核的问题类型和证据。

F1

Missing required section

Major

Candidate B omits the required Limitations section.

F2

Incomplete coverage

Major

Candidate B misses transaction identifier, idempotency, and escalation.

F3

Unsupported claim

Critical

The claim that the refund has settled has no linked evidence record.

F4

Prompt violation

Critical

Candidate B uses the prohibited phrase “refund is guaranteed.”

EXPLICIT WEIGHTS

Keep the ranking logic visible.

  1. Instruction adherence: 35 percent
  2. Factual support: 35 percent
  3. Completeness: 30 percent
  4. Stable tie break: descending score, then candidate identifier

HUMAN REVIEW BOUNDARY

Automation supports judgment. A reviewer owns the decision.

The harness checks explicit text and supplied evidence records. It cannot determine hidden truth, business impact, policy intent, or whether a supplied source is trustworthy. A human reviewer must validate the rubric, evidence quality, edge cases, and final acceptance decision.