Skip to content
All research areas
Research

Natural Language Processing

Language systems that work outside English and outside the demo.

We work on text where the easy assumptions break: low-resource languages, code-switching, domain jargon, and OCR noise. Bengali NLP in particular is under-served, and it is close to home for us.

In production

CMS — reviewer matching & abstract review

Our NLP work powers CMS's topic-based reviewer matching and conflict-of-interest detection — the clearest line from a paper here to a feature in production.

See the product →

What this looks like in practice

  • Information extraction and document understanding
  • Classification, summarisation, and semantic search
  • Low-resource and Bengali language modelling
  • Named entity recognition for specialised domains
  • Corpus construction and annotation pipelines