▲ 62 r/LanguageTechnology+2 crossposts

NLP is growing insanely fast, what will it look like in 2030?

Random thought: NLP in 2010 and NLP in 2020 already felt like two different worlds. The jump was huge.

Now its growing even faster.

So Iam curious how do you think NLP will look in 2030?

What big shifts do you expect? Will it still be mostly scaling transformers or will something completely new take over?

reddit.com
u/CanOk3349 — 5 days ago

Where to focus for NLP Research Scientist Intern roles?

Preparing for NLP Research Scientist Intern roles and overwhelmed by how fast the field moves.

Any advice from people who landed or hire for these roles? What do people waste time on?

Thanks

reddit.com
u/CanOk3349 — 23 days ago
▲ 4 r/u_CanOk3349+2 crossposts

Where to focus for AI Research Scientist Intern roles? Field moves too fast

Trying for AI Research Scientist Intern roles and overwhelmed by how fast everything changes.

What actually matters most right now for interviews/hiring?

Any concrete advice from people who landed or hire for these roles? What do people waste time on?

Thanks.

reddit.com
u/CanOk3349 — 23 days ago
▲ 6 r/LanguageTechnology+1 crossposts

Is making new datasets or fine-tuning still useful in 2026?

Now we have RAG and agentic AI (tools + reasoning). Are making good datasets and fine-tuning old methods? Or are they still better for making AI smarter in special areas?

reddit.com
u/CanOk3349 — 1 month ago

Paper: CLExEval: A Human-in-the-Loop Framework for Qualitative Evaluation of LLM Clinical Reasoning

Link: https://arxiv.org/abs/2606.31608

Summary:
Large language models ace medical exams but struggle with real clinical reasoning. This new paper introduces CLExEval using progressive information masking on rare cases + 5,600 physician annotations.

https://preview.redd.it/x4ema2sitedh1.png?width=822&format=png&auto=webp&s=68e062f82b5bd0128bd41979c26d749f53c29ec0

Key findings:
- Verbosity Bias: GPT-4o-mini accuracy drops from 95% to 32.5% with less info

- Hidden Knowledge Paradox in specialist models

- High Reasoning-Output Mismatch (~69%)

- LLM judges approve a shocking % of clinically wrong outputs

Why it matters: Highlights the evaluation illusion where fluent text masks real failures in high-stakes domains.

What do you think? Is human-in-the-loop evaluation the way forward for clinical AI, or are there better approaches?

(Genuinely interested in discussion)

reddit.com
u/CanOk3349 — 1 month ago
▲ 2 r/ResearchML+1 crossposts

Paper: CLExEval: A Human-in-the-Loop Framework for Qualitative Evaluation of LLM Clinical Reasoning

Link: https://arxiv.org/abs/2606.31608

Summary:

Large language models ace medical exams but struggle with real clinical reasoning. This new paper introduces CLExEval using progressive information masking on rare cases + 5,600 physician annotations.

Key findings:

- Verbosity Bias: GPT-4o-mini accuracy drops from 95% to 32.5% with less info

- Hidden Knowledge Paradox in specialist models

- High Reasoning-Output Mismatch (~69%)

- LLM judges approve a shocking % of clinically wrong outputs

Why it matters: Highlights the "evaluation illusion" where fluent text masks real failures in high-stakes domains.

What do you think? Is human-in-the-loop evaluation the way forward for clinical AI, or are there better approaches?

(Genuinely interested in discussion)

u/CanOk3349 — 1 month ago