Comment-level Topic Drift Analysis in the Reddit Corpus
用Reddit海量评论追踪话题漂移,看社区讨论如何悄然转向。
arXiv:2608.19133v1 Announce Type: new Abstract: We present a novel application of embedding-based dynamic topic modeling techniques to detect and quan…
用Reddit海量评论追踪话题漂移,看社区讨论如何悄然转向。
arXiv:2608.19133v1 Announce Type: new Abstract: We present a novel application of embedding-based dynamic topic modeling techniques to detect and quan…
别再笼统说“AI语言”了,这篇论文用语料库方法论证大模型输出其实带有个人语言习惯特征,视角新颖。
arXiv:2608.06589v1 Announce Type: cross Abstract: While large language model outputs are frequently analysed as a collective super variety termed "AI …
给AI接入圣经语料的MCP服务,让模型轻松查询经文原文。
Article URL: https://nirajagarwal.github.io/bible-mcp/ Comments URL: https://news.ycombinator.com/item?id=49163030 Points: 1 # Comments: 0
大规模指令语料库发布,含近23万文件,可自由用于模型训练与评测。
The TomeVault Instruction Corpus De-identified structural measurements of AI instruction files (CLAUDE.md, AGENTS.md, SKILL.md, .cursorrules and relat…
通过人机协同构建语料库,用大模型简化科学摘要,为AI辅助学术写作提供新方向。
arXiv:2607.25630v1 Announce Type: cross Abstract: Interdisciplinary research is accelerating, yet scientific papers remain difficult to understand out…
为生物学打造的预训练语料库,开启LLM与生命科学深度融合新篇章。
arXiv:2607.08803v1 Announce Type: cross Abstract: The push toward large language models for biology (BioLM) has created a need for training corpora th…
面向社科人文学科,巧用知识图谱与多语言语料让大模型精准适配领域需求。
arXiv:2607.05956v1 Announce Type: new Abstract: The integration of Large Language Models (LLMs) into scientific research workflows, particularly for b…
首个面向学生论文的LLM反馈评估基准,推动自动写作反馈走向可靠落地。
arXiv:2607.00274v1 Announce Type: cross Abstract: Effective writing feedback is among the strongest drivers of student learning, yet producing it at s…
首个阿拉伯语-俄语平行语料库与LLM基准,打破科研语言壁垒,助力可持续发展知识交流。
arXiv:2606.30943v1 Announce Type: new Abstract: Russian and Arabic are among the major languages of scientific communication. Language barriers impede…
为卢森堡语打造的高表现力语音合成语料库,填补低资源语言情感语音研究空白。
arXiv:2606.31947v1 Announce Type: new Abstract: State-of-the-art speech datasets predominantly focus on widely spoken languages, often overlooking low…
语料库驱动AI方法生成的选择题选项可能存在频率偏差,警惕其教育应用中的隐藏陷阱。
arXiv:2602.17377v2 Announce Type: replace Abstract: In recent years, corpus-driven AI methods, such as Large Language Models (LLMs), have seen widespr…
用大语言模型搭建自动化流水线,破解大规模语料库的人工标注瓶颈,高效实现英语语法变异追踪。
arXiv:2510.12306v3 Announce Type: replace Abstract: As natural language corpora expand at an unprecedented rate, manual annotation remains a significa…
粤语-英语语音翻译新基准,600+小时平行语料库开放,句子级对齐助力研究。
arXiv:2306.11252v2 Announce Type: replace-cross Abstract: We introduce HK-LegiCoST, a new three-way parallel corpus of Cantonese-English translations,…
聚焦日本LLM预训练数据中的敏感个人信息检测,为模型训练隐私安全提供新思路
arXiv:2606.12114v1 Announce Type: new Abstract: Sensitive personal information can appear in large-scale pre-training corpora for large language model…
首个科米-亚兹瓦语-俄语平行语料库,为LLM在极低资源语言上的零样本/少样本翻译提供系统评估标准。
arXiv:2606.06420v1 Announce Type: new Abstract: We present the first Komi-Yazva--Russian parallel corpus together with an explicit evaluation protocol…
将Karpathy的LLM教学语料库精心设计为交互式HTML维基,视觉友好、内容系统,适合深度学习大模型知识。
Article URL: https://benzales.github.io/notes/ Comments URL: https://news.ycombinator.com/item?id=48374827 Points: 2 # Comments: 0
提出GrepSeek方法训练LLM搜索代理直接与语料库交互,摆脱传统检索器限制,提升知识密集型任务效率。
arXiv:2605.29307v1 Announce Type: cross Abstract: Large Language Model (LLM) search agents have shown strong promise for knowledge-intensive language …
LLM训练语料库差异会影响科学问答表现吗?一个值得探讨的实验思路
That is, were I to omit all, say, novels and non-fiction from my LLM's training would then scientific questions be better-addressed by that LLM vs ano…
Telenor北欧客户服务自助语料库发布,多语言数据集助力客服AI研究
arXiv:2605.26891v1 Announce Type: new Abstract: This paper presents a multilingual customer service self-help corpus comprising 1,122 manually validat…
介绍EmbGen方法,利用重组语料库革新机器学习教学,探索自适应系统新范式。
arXiv:2605.19394v1 Announce Type: new Abstract: Adapting small instruction-tuned models to specialized domains often relies on supervised fine-tuning …