Natural Memory NM2.1: 记忆路由器分叉、数据集缺陷修复与全轴评测证据
- 引入 MemoryRouterXL 与 v5/v6 流式多线程训练/编码管线 - 修复 prepare_memory_router_dataset 候选池重建缺陷(mega 家族 3568x 加速,输出逐字节相同) - 修复 v5 被破坏的拒答与多跳标签(train 未知样本 319 -> 16319,multi_hop 平均正例 1.00 -> 2.00) - 同存储预算下 V2-128 v6 逐轴 22/22 通过:Top-1 41.12% -> 94.62%,未知拒答 0.00% -> 100.00% - 记录三条被实测推翻的显然优化(logits_to_keep=1 反而慢 55%、XL 容量未带来收益) - 记忆手术跨架构可移植性 14/14,读写关闭时与原生模型逐位相同
This commit is contained in:
@@ -0,0 +1,51 @@
|
||||
"""Exercise the RAG baseline's model-free logic (BM25, prompt, corpus loading).
|
||||
|
||||
Runs without touching the GPU so it can be checked while the memory run owns the
|
||||
card.
|
||||
"""
|
||||
import sys
|
||||
from pathlib import Path
|
||||
|
||||
sys.path.insert(0, str(Path(__file__).resolve().parent.parent))
|
||||
|
||||
from V2_dpskw.strong_rag_locomo import BM25, build_prompt, tokenize
|
||||
from V2_dpskw.eval_end_to_end_memory import build_cases
|
||||
|
||||
# 1. BM25 must rank the turn that actually states the answer first.
|
||||
pool = [
|
||||
"Melanie: I finally finished that painting of a sunrise.",
|
||||
"Caroline: I went to the LGBTQ support group on 7 May 2023.",
|
||||
"Caroline: Work has been really busy this month.",
|
||||
"Melanie: That sounds rough, hope you are okay.",
|
||||
]
|
||||
ranking = BM25(pool).rank("When did Caroline go to the LGBTQ support group?")
|
||||
print("BM25 ranking:", ranking)
|
||||
print("top doc:", pool[ranking[0]])
|
||||
assert ranking[0] == 1, "BM25 failed to rank the answering turn first"
|
||||
|
||||
# 2. A question with no lexical overlap should not crash and should still rank.
|
||||
ranking2 = BM25(pool).rank("What colour was the painting?")
|
||||
print("BM25 (no-overlap) top:", pool[ranking2[0]])
|
||||
|
||||
# 3. Prompt shape.
|
||||
prompt = build_prompt("When did Caroline go?", pool[:2])
|
||||
assert "Question: When did Caroline go?" in prompt and prompt.rstrip().endswith("Answer:")
|
||||
print("prompt ok, length", len(prompt))
|
||||
|
||||
# 4. Tokenizer sanity.
|
||||
print("tokens:", tokenize("I don't know — 7 May 2023!"))
|
||||
assert tokenize("7 May 2023") == ["7", "may", "2023"]
|
||||
|
||||
# 5. Corpus loads with the harness' own loader and the answerable split is right.
|
||||
corpus_path = Path(sys.argv[1] if len(sys.argv) > 1 else r"H:\Memory\V2_dpskw\data\net_locomo\eval.jsonl")
|
||||
cases = build_cases(corpus_path, 40)
|
||||
ans = sum(1 for c in cases if c["answerable"])
|
||||
print(f"corpus: {len(cases)} cases, answerable {ans}, adversarial {len(cases) - ans}")
|
||||
assert ans == 155 and len(cases) == 195, "unexpected answerable split"
|
||||
|
||||
# 6. Every pool has at least a few candidates and no duplicate text.
|
||||
for case in cases:
|
||||
assert len(case["facts"]) >= 4, case["query"]
|
||||
assert len(set(case["facts"])) == len(case["facts"]), "duplicate candidate text"
|
||||
print("all pools >= 4 candidates, no duplicates")
|
||||
print("\nOK")
|
||||
Reference in New Issue
Block a user