中文简介
TRuST 是一个俄语网页检索基准,包含 324 道需要组合多条证据的短答案问题,用于评估语言模型和搜索智能体的检索能力。上游明确说明来源来自抓取时公开的网页且不主张原始材料所有权,因此用户需自行审查来源、网站条款、准确性和研究使用边界。
上游模型卡 / 数据集卡
TRuST 是一个俄语网页检索基准,包含 324 道需要组合多条证据的短答案问题,用于评估语言模型和搜索智能体的检索能力。上游明确说明来源来自抓取时公开的网页且不主张原始材料所有权,因此用户需自行审查来源、网站条款、准确性和研究使用边界。
TRuST: T-Tech Russian Search Test 🚨 **TRuST was built from sources collected from the open web that were publicly accessible at the time of crawling. We do not claim ownership of the original source materials and do not endorse, verify, or take responsibility for the accuracy, completeness, legality, or safety of the information contained in those sources. It is intended for research and development purposes only. Users are fully responsible for inspecting the data, respecting applicable licenses and website terms, and ensuring that any downstream use complies with relevant ethical, legal, and safety requirements.** 📚 Description TRuST** is a Russian BrowseComp-Plus-like web-search benchmark designed to evaluate the retrieval abilities of language models and search agents. It contains **324** hard compositional short-answer questions over the Russian web. TRuST was created manually by human annotators: each question was written, verified, and linked to gold supporting documents that contain the evidence required to derive the final answer. The benchmark is designed to measure how well a model can find the right evidence in a fixed search index, especially in Russian-language and Runet-specific information-seeking scenarios. 📝 Benchmark Summary Benchmark covers a broad range of Russian-language information-seeking tasks across 8 topic categories: Business & Technology Sports News & Politics Science & Academic Publications Media & Art Regulation & Law Geography History & Archives The benchmark also includes 5 types of search challenges, inspired by the SealQA taxonomy of retrieval-oriented question types, that reflect different retrieval failure modes: **Multihop** — tasks where the answer cannot be found in a single source directly. The retriever has to connect evidence across several documents, entities, or intermediate facts. **Structured evidence** — tasks that require extracting information from structured or semi-structured sources, such as tables, lists, regist
上游文件元数据
.gitattributes2.52 KBbenchmark/data/corpus/data_part_0.parquet275.77 MBbenchmark/data/corpus/data_part_1.parquet277.90 MBbenchmark/data/corpus/data_part_2.parquet275.27 MBbenchmark/data/corpus/data_part_3.parquet283.69 MBbenchmark/data/topics-qrels/qrel_evidence.txt27.95 KBbenchmark/data/topics-qrels/queries.encrypted.tsv290.55 KBbenchmark/data/trust_encrypted.jsonl79.91 MBbenchmark/indexes/qwen3-embedding-8b/corpus.shard1_of_4.pkl10.41 GBbenchmark/indexes/qwen3-embedding-8b/corpus.shard2_of_4.pkl10.41 GBbenchmark/indexes/qwen3-embedding-8b/corpus.shard3_of_4.pkl10.41 GBbenchmark/indexes/qwen3-embedding-8b/corpus.shard4_of_4.pkl10.41 GB
本页面为橙子AI科技的中文整理与服务说明,不代表资源作者或平台官方页面。实际许可、访问和使用条件以上游原文为准。