Missing required section
Major
Candidate B omits the required Limitations section.
FICTIONAL LLM EVALUATION SAMPLE
A deterministic Python harness that ranks two fictional AI answers against visible requirements, evidence records, and prohibited claims. It makes no model calls and produces the same structured report from the same input.
TYPED FINDINGS
中文摘要:该虚构样例用明确权重评估指令遵循、事实支持与完整性,并输出可复核的问题类型和证据。
Major
Candidate B omits the required Limitations section.
Major
Candidate B misses transaction identifier, idempotency, and escalation.
Critical
The claim that the refund has settled has no linked evidence record.
Critical
Candidate B uses the prohibited phrase “refund is guaranteed.”
EXPLICIT WEIGHTS
HUMAN REVIEW BOUNDARY
The harness checks explicit text and supplied evidence records. It cannot determine hidden truth, business impact, policy intent, or whether a supplied source is trustworthy. A human reviewer must validate the rubric, evidence quality, edge cases, and final acceptance decision.