How Reproducible Are Evaluation Conclusions? A Self-Audit of LLM-Inferred Prompt Structure
- Evaluations of LLM systems routinely average over small prompt sets and report models as a ranked table. We ask how much
1媒体
過去7日に報じた媒体数を数えています。同じ媒体が何度書いても1件。公式の一次情報は媒体数に含めず、別に数えます。直近24時間の初報 1媒体、その前の24時間 0媒体。見出しが一致する転載候補 0件は媒体数から除外しました。初報は13時間前。