Publications

This section contains publications published by ISI over the past several decades.

2021

Modeling" newsworthiness" for lead-generation across corpora

Alexander Spangher, Nanyun Peng, Jonathan May, Emilio Ferrara
arXiv preprint arXiv:2104.09653,  2021

Cross-attention is all you need: Adapting pretrained transformers for machine translation

Mozhdeh Gheini, Xiang Ren, Jonathan May
arXiv preprint arXiv:2104.08771,  2021

Macro-average: rare types are important too

Thamme Gowda, Weiqiu You, Constantine Lignos, Jonathan May
arXiv preprint arXiv:2104.05700,  2021

Many-to-English machine translation tools, data, and pretrained models

Thamme Gowda, Zhao Zhang, Chris A Mattmann, Jonathan May
arXiv preprint arXiv:2104.00290,  2021

CaSiNo: A corpus of campsite negotiation dialogues for automatic negotiation systems

Kushal Chawla, Jaysa Ramirez, Rene Clever, Gale Lucas, Jonathan May, Jonathan Gratch
arXiv preprint arXiv:2103.15721,  2021

Multitask learning for class-imbalanced discourse classification

Alexander Spangher, Jonathan May, Sz-rung Shiang, Lingjia Deng
arXiv preprint arXiv:2101.00389,  2021

On the strengths of cross-attention in pretrained transformers for machine translation

Mozhdeh Gheini, Xiang Ren, Jonathan May
arXiv preprint arXiv:2104.08771,  2021

CheckDST: Measuring real-world generalization of dialogue state tracking performance

Hyundong Cho, Chinnadhurai Sankar, Christopher Lin, Kaushik Ram Sadagopan, Shahin Shayandeh, Asli Celikyilmaz, Jonathan May, Ahmad Beirami
arXiv preprint arXiv:2112.08321,  2021

\textit {StateCensusLaws. org}: A Web Application for Consuming and Annotating Legal Discourse Learning.

Alexander Spangher, Jonathan May
arXiv preprint arXiv:2104.10263,  2021

Know thy strengths: Comprehensive dialogue state tracking diagnostics

Hyundong Cho, Chinnadhurai Sankar, Christopher Lin, Kaushik Ram Sadagopan, Shahin Shayandeh, Asli Celikyilmaz, Jonathan May, Ahmad Beirami
arXiv preprint arXiv:2112.08321,  2021

Good Examples Make A Faster Learner: Simple Demonstration-based Learning for Low-resource NER

Dong-Ho Lee, Akshen Kadakia, Kangmin Tan, Mahak Agarwal, Xinyu Feng, Takashi Shibuya, Ryosuke Mitani, Toshiyuki Sekiya, Jay Pujara, Xiang Ren
ACL 2022,  2021

Improving Text Auto-Completion with Next Phrase Prediction

Dong-Ho Lee, Zhiqiang Hu, Roy Ka-Wei Lee
EMNLP 2021 Findings,  2021

Speaker turn modeling for dialogue act classification

Zihao He, Leili Tavabi, Kristina Lerman, Mohammad Soleymani
arXiv preprint arXiv:2109.05056,  2021

Pattern discovery in time series with byte pair encoding

Nazgol Tavabi, Kristina Lerman
arXiv preprint arXiv:2106.00614,  2021

Detecting polarized topics using partisanship-aware contextualized topic embeddings

Zihao He, Negar Mokhberian, António Câmara, Andrés Abeliuk, Kristina Lerman
arXiv preprint arXiv:2104.07814,  2021

Gender disparity in the authorship of biomedical research publications during the COVID-19 pandemic: retrospective observational study

Goran Muric, Kristina Lerman, Emilio Ferrara
Journal of medical Internet research 23 (4), e25379,  2021

Follow the leader: Documents on the leading edge of semantic change get more citations

Sandeep Soni, Kristina Lerman, Jacob Eisenstein
Journal of the Association for Information Science and Technology 72 (4 …,  2021

Emergence of structural inequalities in scientific citation networks

Buddhika Nettasinghe, Nazanin Alipourfard, Vikram Krishnamurthy, Kristina Lerman
arXiv preprint arXiv:2103.10944,  2021

Explaining classification performance and bias via network structure and sampling

Lisette Espín‑Noboa, Fariba Karimi, Bruno Ribeiro, Kristina Lerman, Claudia Wagner

A Grounded Approach to Modeling Generic Knowledge Acquisition

Deniz Beser, Joe Cecil, Marjorie Freedman, Jacob Lichtefeld, Mitch Marcus, Sarah Payne, Charles Yang
arXiv preprint arXiv:2105.03207,  2021