SPLIT拆问题
Conduct + stance + outside evidence把行为、立场和外部证据分开正在加载页面
正在加载页面
MPHIL RESEARCH DISCUSSION · 研究方案讨论
两种标签来源,接受同一场公平考试。
Same model recipe. Different training labels.同一模型配方,不同训练标签。
RESEARCH QUESTION · 研究问题
Can we find risky trading promotion in a new channel—without wrongly flagging victims and warnings?能否在新频道里找到高风险交易推广,同时避免误伤受害者叙述和警示信息?SPLIT拆问题
Conduct + stance + outside evidence把行为、立场和外部证据分开TRAIN造信号
Human labels vs auditable weak labels比较人工标签与可审计弱标签TEST考模型
One permitted, held-out channel用一个获许可的留出频道考试OUTPUT · two model report cards; any later use is a human-review queue, not a verdict产物 · 两张模型成绩单;未来若使用,也只输出人工复核排序,不作认定
01 · SPLIT THE PROBLEM
CONTROLLED COMPARISON受控对照
“稳赚” + “名额快满,现在入金”“这单稳赚,名额快满了。现在入金,我拉你进群。”
Issues and supports it直接发出并支持这项承诺
“他当时说‘这单稳赚,名额快满了,现在入金’,我照做后亏了钱。”
Attributes it to someone else把话术归给他人并报告受损
“有人说‘这单稳赚,名额快满了,现在入金’,不要相信,先核验监管信息。”
Quotes and rejects it引述后明确否定并提醒别人
Which risky proposition appears?消息里出现了哪个风险命题?
How does this speaker relate to it?当前说话者是在背书,还是引述后报告/警示?
What separate source exists outside?消息之外另有哪些可追溯资料?
ILLUSTRATIVE REWRITES · 示意改写例句,不使用原始消息
02 · MAKE TRAINING SIGNALS
SUPERVISED BASELINE · 监督基线
WEAK SUPERVISION · 弱监督
SAME RECIPE · 同一种配方
Same model recipe, trained twice同一种模型配方,分别训练两次SEALED EXAM · 封存考卷
Humans label messages; software grades models人给消息标答案;程序给模型算成绩The main difference is where the training labels come from.核心只比较:训练标签从哪里来。
TF-IDF + logistic regression / TF-IDF+逻辑回归
Different messages · never used for training · humans label messages and software computes both model grades另一批消息 · 不参与任何训练 · 人给消息标答案,程序给两个模型算成绩
Human labels appear twice for two different jobs: training labels teach Route 1; test labels teach neither model and become the shared answer key.人工标签出现两次,但工作不同:训练标签只教路线1;考试标签不教任何模型,只形成一份共同答案册。
Only the training-label count is matched. Rule design and maintenance effort are reported separately.只匹配训练标签数量;规则设计与维护工时另行报告。
03 · TEST THE MODEL
GATE 0 · WRITTEN PERMISSION FIRST第0关 · 先取得书面许可
Only permitted channel-time windows may enter the experiment.只有获许可的社群与时间窗口才能进入实验。Ethics/governance, consent or waiver where needed, source/platform authority, privacy, security and retention are separate gates.伦理/治理、需要时的同意或豁免、来源/平台权限、隐私、安全与留存是相互独立的门。
CURRENTLY PLANNED · no permitted Chinese multi-channel stream yet当前仍是方案 · 尚无获许可的中文多频道消息流
LOCK THE CHANNEL锁频道
Hold out one permitted channel before training.训练前封存一个获许可的频道。LOCK REPOSTS锁转发家族
Keep each repost family on only one side.同一转发家族只留在边界一侧。GRADE ONCE同一答案册
One human answer key; software makes two report cards.人工制作一份答案;程序计算两张成绩单。PRIMARY主结果
Ranking quality目标消息是否排在前面Average Precision · 平均精确率SAFETY安全线
Protected cases wrongly flagged受害者与警示是否被误伤Harmful false alerts · 有害误报OUTPUT输出
A ranked queue for human review交给人工优先复核的排序Not a verdict · 不作认定Design stage · no Chinese results yet · ranking for human review only方案阶段 · 尚无中文实证结果 · 只做人审排序
One copied template lands on both sides. The model may remember wording, so the score can look too good.同一转发模板同时落入训练和考试。模型可能只记住话术,让分数看起来虚高。
‘Unseen’ means unseen by model development—not unapproved, secret or unmanaged.“未见”是模型开发过程未见,不是未经许可、秘密进入或无人管理。
Channel holdout tests transfer to a different community. Repost-family control prevents memorising the same template.整频道留出检验换一个社群还能不能用;转发家族控制防止模型背同一套话术。Later-time is secondary: another permitted batch after a pre-set date, with the same cross-boundary repost purge. It is labelled and scored offline.后来时间是次要复测:预设日期后到达、同样获许可的批次,也执行跨边界重复家族剔除;它离线标注、离线评分。
Average Precision / 平均精确率
Harmful false alerts / 有害误报
Keep a count and report coverage loss.保留数量记录,并报告因此损失了多少覆盖率。
The question · 开题