I am a computer science researcher and professor specializing in human-AI interaction at Northeastern University's Khoury College of Computer Sciences. My research asks a single question in several forms: how do we know whether a language model is doing what we intended, and how does that question change when the model acts rather than merely responds? Once a system selects tools, issues calls, and consumes its own outputs as inputs, the unit of evaluation shifts from a response to a trajectory, and most of our inherited measurement apparatus stops applying cleanly. My work develops the evaluation methods, system architectures, and human-facing interfaces that this shift demands.
In my research, I study how model behavior varies under conditions that evaluation suites typically hold fixed. In recent work I examined how personalization cues — user context disclosed in the prompt — shift refusal behavior and jailbreak task completion across multiple deployed systems. The finding I take to be robust is that safety behavior is not a stable model property but a joint function of model and context; the stronger claim, that specific cue types produce predictable shifts across architectures, remains a hypothesis my current data underdetermines.
This line extends naturally to agentic settings, where failure is distributed across steps rather than localized in an output. I am interested in capability taxonomies that survive contact with multi-turn tool use, in step-level rather than trajectory-level attribution of failure, and in adversarial robustness for benchmarks whose environments are themselves generated. A parallel strand considers how such evaluations should be governed — in particular, tiered-access models for benchmarks whose diagnostic value depends on not being fully public.
A second strand concerns how small language models acquire tool use, and what this implies for the governability of agentic systems. The prevailing SLM-first argument is economic: most agent invocations are narrow, repetitive, and well within the reach of a fine-tuned small model. My interest is that the same property makes these systems more auditable. A heterogeneous system in which a planner delegates to specialized small workers exposes a decomposition that a monolithic model does not, and each component admits targeted evaluation. I work on trajectory distillation and training-free memory transfer as routes to competent small tool-users, on the failure modes that distinguish them from large models — tool selection, argument construction, and knowing when no tool applies — and on what on-device inference makes newly possible for domains where data cannot leave the institution.
My research contributions span human-AI interaction, AI-driven user experiences, immersive technologies (XR), and accessibility, with publications in leading conferences and journals. My work has been featured in venues such as ACM CHI, ACM UIST, IEEE AIVR, Transactions on Visualization and Computer Graphics, and Journal of Human-Computer Interaction.
For a full list of my publications, please visit my Google Scholar page
I have extensive experience teaching and mentoring AI developers and researchers. Currently, I'm teaching courses on Agentic AI and AI Evals at Northeastern University.
I am humbled by the scholarly interest in using the Nomophobia Questionnaire (NMP-Q) in research studies. Please feel free to use the NMP-Q in your research studies without seeking permission. You can access the relevant articles under Publications. You can also download the questionnaire, along with the scoring guide, below.