Skip to content
Sandip Chapagain

Research

Research interests

Language technology for Nepali and multilingual settings: how to represent, search and retrieve text across languages, and how large language models behave on low-resource languages.

Areas of interest

  • Large language models

    How LLMs behave on low-resource languages, and how to use them reliably in real applications.

  • Multilingual embeddings

    Representing text from different languages in a shared vector space so meaning can be compared across languages.

  • Nepali language processing

    Tools, datasets and models for Nepali, a language that is still under-served by mainstream NLP.

  • Natural language processing

    Core NLP problems such as tokenisation, classification and retrieval, with a practical engineering focus.

  • Semantic search

    Search that matches meaning rather than exact keywords, built on embeddings and vector retrieval.

  • Multilingual information retrieval

    Finding relevant documents when the query and the documents may be in different languages, such as Nepali and English.

Publications

No publications yet. Papers, preprints and datasets will be listed here with their abstracts, DOIs, code and citation information as they are released.