# 中文大语言模型评测第三期

**URL:** <http://forum.beginner.center/t/topic/2523>\
**Category:** 🛠工具与编程\
**Tags:** 评测\
**Created:** [2025 年12 月 20 日 03:22 UTC](http://forum.beginner.center/t/topic/2523 "2025-12-20T03:22:10Z")\
**Posts on this page:** 1\
**Page:** 1

<div class="post-metadata">

**Author:** ![doggie](http://forum.beginner.center/user_avatar/forum.beginner.center/doggie/32/550_2.png) [@doggie](http://forum.beginner.center/u/doggie)\
**Post date:** [2025 年12 月 20 日 03:22 UTC](http://forum.beginner.center/t/topic/2523/1 "2025-12-20T03:22:10Z")

</div>

> **[GitHub - llmeval/LLMEval-Fair: \[ACL 2026\] A large-scale longitudinal study on...](https://github.com/llmeval/LLMEval-Fair)**
>
> \[ACL 2026\] A large-scale longitudinal study on robust and fair evaluation of LLMs — 200K+ generative questions across 13 disciplines

> **[LLMEval — Comprehensive LLM Evaluation](https://llmeval.com/)**
>
> LLMEval is a research series dedicated to building comprehensive, fair, and robust evaluation frameworks for large language models.

> **[LLMEval-Fair: A Large-Scale Longitudinal Study on Robust and Fair Evaluation...](https://arxiv.org/abs/2508.05452)**
>
> Existing evaluation of Large Language Models (LLMs) on static benchmarks is vulnerable to data contamination and leaderboard overfitting, critical issues that obscure true model capabilities. To address this, we introduce LLMEval-Fair, a framework...
