Publications

For research publications from the ARIA Research Lab, please visit the ARIA Lab publications page. You can also visit Dr. Arya’s Google Scholar profile.

Dissertation

  1. Arya, Shivvrat
    ProQuest Dissertations and Theses, 2025
    PDF
    Abstract
    Probabilistic Models (PMs) provide a powerful framework for representing large multivariate probability distributions, but as their expressiveness increases, inference tasks become computationally intractable. In the era of big data, the demand for complex models capable of effectively representing and reasoning over large-scale data has increased significantly. While models like probabilistic circuits allow tractable inference for specific query types, handling a broader range of queries remains challenging. Exact inference in PMs is generally NP-hard, and approximate methods, though viable, often involve trade-offs between accuracy and efficiency, leading to unreliable estimates. Achieving near-optimal solutions typically demands substantial computational resources, posing a barrier to scalability.This dissertation presents novel algorithms for key inference tasks in PMs, including Most Probable Explanation (MPE), Constrained Most Probable Explanation (CMPE), and Marginal Maximum A Posteriori (MMAP) inference. Our methods generalize across different Probabilistic Models (PMs), such as Probabilistic Graphical Models (PGMs), Probabilistic Circuits (PCs), and Neural Autoregressive Models (NAMs), enabling efficient and scalable probabilistic reasoning. Additionally, we propose a neurosymbolic framework that integrates Dependency Networks (DNs) with deep learning architectures to enhance multi-label classification. Specifically, this dissertation makes the following contributions:We develop a self-supervised learning framework for neural network-based solvers to answer predefined MPE and MMAP queries over PCs. Our approach introduces a scalable loss function with a computation cost that scales linearly with the size of the PC, enabling the training of neural solvers that match or surpass state-of-the-art solvers in solution quality while reducing inference time from seconds to microseconds.We extend inference capabilities beyond predefined queries, allowing neural solvers to answer arbitrary MPE queries over PCs, PGMs, and NAMs. To achieve this, we introduce a novel dual-network learning framework and an enhanced inference scheme that updates neural network parameters at test time, leading to improved solution quality.We eliminate the need for costly inference-time optimization in neural solvers for arbitrary MPE queries over PGMs by enhancing query encoding, which improves solution quality through richer embeddings. Additionally, we introduce two novel methods for discretizing continuous neural network outputs, further enhancing solution quality.We extend neural solvers to constrained inference, specifically the Constrained Most Probable Explanation (CMPE) task. We develop a self-supervised CMPE solver with a loss function satisfying the consistent loss property, ensuring alignment with the optimal CMPE solution—unlike existing loss functions for constrained optimization, which lack this guarantee.We propose advanced inference schemes for MPE in neurosymbolic models for multi-label classification, specifically improving inference in Deep Dependency Networks (DDNs). While DDNs offer efficient training and an intuitive loss function for multi-label classification, they traditionally rely on Gibbs sampling, which limits inference accuracy and efficiency. To address this limitation, we introduce novel inference techniques based on local search and integer linear programming (ILP), facilitating more accurate and efficient computation of the most probable label assignments.

Publications

  1. Shahriari, Reza and Hashky, Amal and Arya, Shivvrat and Audino, Tyler and Ragan, Eric D. and Gogate, Vibhav and Ruiz, Jaime
    ACM Transactions on Interactive Intelligent Systems, 2026
    DOI
    Abstract
    Human-in-the-loop methods leverage human feedback to enhance machine learning and AI. Manual review of outputs can correct errors, identify model weaknesses, or expand labels to broaden model capabilities. Feedback collection methods range from simple flagging of outputs as correct or incorrect to more complex feature-level adjustments or natural language interpretations. This article presents a user study evaluating changes in user performance over time and explores the tradeoff between feedback quality and human effort. We compare four interactive input methods for reviewing and correcting outcomes in object detection and activity recognition in videos. Our findings indicate that while some complex input methods, such as free-text, require more time, the quality and impact of their feedback on model accuracy often surpass those of simpler methods that require less effort. However, more effort does not always lead to better-quality feedback, especially when aiming to improve the model. Our VLM experiments show that the most accurate models were trained using detailed natural language feedback or precise word-level corrections, while simple yes/no judgments also led to solid performance at a much lower annotation cost.
  2. Vyas, Akshay and Arya, Shivvrat and Malhotra, Brij and Rahman, Tahrima and Gogate, Vibhav Giridhar and Ruozzi, Nicholas
    Advances in Neural Information Processing Systems, 2026
    PDF
    Spotlight presentation (top 1% of papers)
  3. Chen, Yeqing and Qi, Cong and Fang, Hanzhang and Luan, Feiyang and Zhang, Zhirong and Arya, Shivvrat and Wei, Zhi
    bioRxiv, 2026
    DOI Code
    Abstract
    Single-cell RNA sequencing provides a powerful view of cellular heterogeneity, but its sparsity and dropout noise remain major obstacles for recovering biologically meaningful gene expression programs and for downstream analyses that depend on reliable expression measurements. Ligand–receptor-based cell–cell communication inference is such analysis, missing ligand or receptor expression can cause substantial false negatives in sparse single-cell data. Here, we present CoLa-VAE, a cell–cell communication-aware variational autoencoder that jointly learns latent representations and denoised expression profiles by incorporating ligand–receptor-derived communication topology through dynamic graph Laplacian regularization. Rather than treating denoising as a secondary output of representation learning, CoLa-VAE uses denoised expression to iteratively refine communication estimates and uses the resulting communication structure to guide both latent organization and expression reconstruction. In addition to improving latent space organization and producing robust denoised expression matrices, CoLa-VAE-denoised matrices also improved downstream biological analyses, including the detection of robust differential cell–cell communication programs, mitigation of batch-associated variation and enhanced spatial transcriptomic deconvolution when spatially constrained communication structure was incorporated. Together, these results establish CoLa-VAE as a communication-guided denoising and representation learning framework that recovers biologically meaningful expression signals from sparse single-cell and spatial transcriptomic data, enabling more sensitive and reliable downstream analysis.Competing Interest StatementThe authors have declared no competing interest.
  4. Malhotra, Brij and Arya, Shivvrat and Rahman, Tahrima and Gogate, Vibhav Giridhar
    arXiv preprint arXiv:2602.01475, 2026
    PDF arXiv
  5. Arya, Shivvrat and Ghosh, Smita and Maruyama, Bryan and Srinivasan, Venkatesh
    Proceedings of the 34th ACM International Conference on Information and Knowledge Management, 2025
    PDF DOI
    Oral presentation
  6. Malhotra, Brij and Arya, Shivvrat and Rahman, Tahrima and Gogate, Vibhav
    Advances in Neural Information Processing Systems, 2025
    PDF Code
    Abstract
    We introduce learning to condition (L2C), a scalable, data-driven framework for accelerating Most Probable Explanation (MPE) inference in Probabilistic Graphical Models (PGMs), a fundamentally intractable problem. L2C trains a neural network to score variable-value assignments based on their utility for conditioning, given observed evidence. To facilitate supervised learning, we develop a scalable data generation pipeline that extracts training signals from the search traces of existing MPE solvers. The trained network serves as a heuristic that integrates with search algorithms, acting as a conditioning strategy prior to exact inference or as a branching and node selection policy within branch-and-bound solvers. We evaluate L2C on challenging MPE queries involving high-treewidth PGMs. Experiments show that our learned heuristic significantly reduces the search space while maintaining or improving solution quality over state-of-the-art methods.
  7. Arya, Shivvrat and Rahman, Tahrima and Gogate, Vibhav Giridhar
    Proceedings of The 28th International Conference on Artificial Intelligence and Statistics, 2025
    PDF Code Library
    Abstract
    Our paper builds on the recent trend of using neural networks trained with self-supervised or supervised learning to solve the Most Probable Explanation (MPE) task in discrete graphical models. At inference time, these networks take an evidence assignment as input and generate the most likely assignment for the remaining variables via a single forward pass. We address two key limitations of existing approaches: (1) the inability to fully exploit the graphical model’s structure and parameters, and (2) the suboptimal discretization of continuous neural network outputs. Our approach embeds model structure and parameters into a more expressive feature representation, significantly improving  performance. Existing methods rely on standard thresholding, which often yields suboptimal results due to the non-convexity of the loss function. We introduce two methods to overcome discretization challenges: (1) an external oracle-based approach that infers uncertain variables using additional evidence from confidently predicted ones, and (2) a technique that identifies and selects the highest-scoring discrete solutions near the continuous output. Experimental results on various probabilistic models demonstrate the effectiveness and scalability of our approach, highlighting its practical impact.
  8. Predictive Task Guidance with Artificial Intelligence in Augmented Reality
    Rheault, Benjamin and Arya, Shivvrat and Vyas, Akshay and Wang, Jikai and Peddi, Rohith and Benda, Brett and Gogate, Vibhav and Ruozzi, Nicholas and Xiang, Yu and Ragan, Eric D
    IEEE Virtual Reality (IEEE VR), 2024
    PDF
  9. Arya, Shivvrat and Rahman, Tahrima and Gogate, Vibhav Giridhar
    The 7th Workshop on Tractable Probabilistic Modeling (TPM), 2024
    PDF Library
  10. Arya, Shivvrat and Rahman, Tahrima and Gogate, Vibhav Giridhar
    The 7th Workshop on Tractable Probabilistic Modeling (TPM), 2024
    PDF Library
    Best Paper Award
    Certificate
    Abstract
    We propose a novel neural networks based approach to efficiently answer arbitrary Most Probable Explanation (MPE) queries—a well-known NP-hard task—in large probabilistic models such as Bayesian and Markov networks, probabilistic circuits, and neural auto-regressive models. By arbitrary MPE queries, we mean that there is no predefined partition of variables into evidence and non-evidence variables. The key idea is to distill all MPE queries over a given probabilistic model into a neural network and then use the latter for answering queries, eliminating the need for time-consuming inference algorithms that operate directly on the probabilistic model. We improve upon this idea by incorporating inference-time optimization with self-supervised loss to iteratively improve the solutions and employ a teacher-student framework that provides a better initial network, which in turn, helps reduce the number of inference-time optimization steps. The teacher network utilizes a self-supervised loss function optimized for getting the exact MPE solution, while the student network learns from the teacher’s near-optimal outputs through supervised loss. We demonstrate the efficacy and scalability of our approach on various datasets and a broad class of probabilistic models, showcasing its practical effectiveness.
  11. Arya, Shivvrat and Rahman, Tahrima and Gogate, Vibhav
    Proceedings of The 27th International Conference on Artificial Intelligence and Statistics (AISTATS), 2024
    PDF Code Library
    Abstract
    We propose a self-supervised learning approach for solving the following constrained optimization task in log-linear models or Markov networks. Let f and g be two log-linear models defined over the sets X and Y of random variables respectively. Given an assignment x to all variables in X (evidence) and a real number q, the constrained most-probable explanation (CMPE) task seeks to find an assignment y to all variables in Y such that f(x,y) is maximized and g(x,y)≤q. In our proposed self-supervised approach, given assignments x to X (data), we train a deep neural network that learns to output near-optimal solutions to the CMPE problem without requiring access to any pre-computed solutions. The key idea in our approach is to use first principles and approximate inference methods for CMPE to derive novel loss functions that seek to push infeasible solutions towards feasible ones and feasible solutions towards optimal ones. We analyze the properties of our proposed method and experimentally demonstrate its efficacy on several benchmark problems.
  12. Arya, Shivvrat and Rahman, Tahrima and Gogate, Vibhav
    Proceedings of the AAAI Conference on Artificial Intelligence (Oral), 2024
    PDF DOI Library
    Oral presentation (top 3% of papers)
    Abstract
    Probabilistic circuits (PCs) such as sum-product networks efficiently represent large multi-variate probability distributions. They are preferred in practice over other probabilistic representations, such as Bayesian and Markov networks, because PCs can solve marginal inference (MAR) tasks in time that scales linearly in the size of the network. Unfortunately, the most probable explanation (MPE) task and its generalization, the marginal maximum-a-posteriori (MMAP) inference task remain NP-hard in these models. Inspired by the recent work on using neural networks for generating near-optimal solutions to optimization problems such as integer linear programming, we propose an approach that uses neural networks to approximate MMAP inference in PCs. The key idea in our approach is to approximate the cost of an assignment to the query variables using a continuous multilinear function and then use the latter as a loss function. The two main benefits of our new method are that it is self-supervised, and after the neural network is learned, it requires only linear time to output a solution. We evaluate our new approach on several benchmark datasets and show that it outperforms three competing linear time approximations: max-product inference, max-marginal inference, and sequential estimation, which are used in practice to solve MMAP tasks in PCs.
  13. CaptainCook4D: a dataset for understanding errors in procedural activities
    Peddi, Rohith and Arya, Shivvrat and Challa, Bharath and Pallapothula, Likhitha and Vyas, Akshay and Gouripeddi, Bhavya and Zhang, Qifan and Wang, Jikai and Komaragiri, Vasundhara and Ragan, Eric and Ruozzi, Nicholas and Xiang, Yu and Gogate, Vibhav
    Proceedings of the 38th International Conference on Neural Information Processing Systems, 2024
    PDF Code Website
    Abstract
    Following step-by-step procedures is an essential component of various activities carried out by individuals in their daily lives. These procedures serve as a guiding framework that helps to achieve goals efficiently, whether it is assembling furniture or preparing a recipe. However, the complexity and duration of procedural activities inherently increase the likelihood of making errors. Understanding such procedural activities from a sequence of frames is a challenging task that demands an accurate interpretation of visual information and the ability to reason about the structure of the activity. To this end, we collect a new egocentric 4D dataset CaptainCook4D comprising 384 recordings (94.5 hours) of people performing recipes in real kitchen environments. This dataset consists of two distinct types of activities: one in which participants adhere to the provided recipe instructions and another in which they deviate and induce errors. We provide 5.3K step annotations and 10K finegrained action annotations and benchmark the dataset for the following tasks: error recognition, multi-step localization and procedure learning. https://captaincook4d.github.io/captain-cook/
  14. Arya, Shivvrat and Rahman, Tahrima and Gogate, Vibhav
    Advances in Neural Information Processing Systems, 2024
    PDF Library
    Spotlight presentation (top 2% of papers)
    Abstract
    We propose a novel neural networks based approach to efficiently answer arbitrary Most Probable Explanation (MPE) queries—a well-known NP-hard task—in large probabilistic models such as Bayesian and Markov networks, probabilistic circuits, and neural auto-regressive models. By arbitrary MPE queries, we mean that there is no predefined partition of variables into evidence and non-evidence variables. The key idea is to distill all MPE queries over a given probabilistic model into a neural network and then use the latter for answering queries, eliminating the need for time-consuming inference algorithms that operate directly on the probabilistic model. We improve upon this idea by incorporating inference-time optimization with self-supervised loss to iteratively improve the solutions and employ a teacher-student framework that provides a better initial network, which in turn, helps reduce the number of inference-time optimization steps. The teacher network utilizes a self-supervised loss function optimized for getting the exact MPE solution, while the student network learns from the teacher’s near-optimal outputs through supervised loss. We demonstrate the efficacy and scalability of our approach on various datasets and a broad class of probabilistic models, showcasing its practical effectiveness.
  15. Arya, Shivvrat and Xiang, Yu and Gogate, Vibhav
    Proceedings of The 27th International Conference on Artificial Intelligence and Statistics, 2024
    PDF Code
    Abstract
    We present a unified framework called deep dependency networks (DDNs) that combines dependency networks and deep learning architectures for multi-label classification, with a particular emphasis on image and video data. The primary advantage of dependency networks is their ease of training, in contrast to other probabilistic graphical models like Markov networks. In particular, when combined with deep learning architectures, they provide an intuitive, easy-to-use loss function for multi-label classification. A drawback of DDNs compared to Markov networks is their lack of advanced inference schemes, necessitating the use of Gibbs sampling. To address this challenge, we propose novel inference schemes based on local search and integer linear programming for computing the most likely assignment to the labels given observations. We evaluate our novel methods on three video datasets (Charades, TACoS, Wetlab) and three image datasets (MS-COCO, PASCAL VOC, NUS-WIDE), comparing their performance with (a) basic neural architectures and (b) neural architectures combined with Markov networks equipped with advanced inference and learning techniques. Our results demonstrate the superiority of our new DDN methods over the two competing approaches.
  16. Roy*, Chiradeep and Nourani*, Mahsan and Arya*, Shivvrat and Shanbhag, Mahesh and Rahman, Tahrima and Ragan, Eric D. and Ruozzi, Nicholas and Gogate, Vibhav (*equal contribution)
    ACM Transactions on Interactive Intelligent Systems (TiiS), 2023
    DOI
    Abstract
    We consider the following video activity recognition (VAR) task: given a video, infer the set of activities being performed in the video and assign each frame to an activity. Although VAR can be solved accurately using existing deep learning techniques, deep networks are neither interpretable nor explainable and as a result their use is problematic in high stakes decision-making applications (e.g., in healthcare, experimental Biology, aviation, law, etc.). In such applications, failure may lead to disastrous consequences and therefore it is necessary that the user is able to either understand the inner workings of the model or probe it to understand its reasoning patterns for a given decision. We address these limitations of deep networks by proposing a new approach that feeds the output of a deep model into a tractable, interpretable probabilistic model called a dynamic conditional cutset network that is defined over the explanatory and output variables and then performing joint inference over the combined model. The two key benefits of using cutset networks are: (a) they explicitly model the relationship between the output and explanatory variables and as a result the combined model is likely to be more accurate than the vanilla deep model and (b) they can answer reasoning queries in polynomial time and as a result they can derive meaningful explanations by efficiently answering explanation queries. We demonstrate the efficacy of our approach on two datasets, Textually Annotated Cooking Scenes (TACoS), and wet lab, using conventional evaluation measures such as the Jaccard Index and Hamming Loss, as well as a human-subjects study.
  17. Arya, Shivvrat and Xiang, Yu and Gogate, Vibhav
    arXiv preprint arXiv:2302.00633, 2023
    PDF arXiv
    Abstract
    We propose a simple approach which combines the strengths of probabilistic graphical models and deep learning architectures for solving the multi-label classification task, focusing specifically on image and video data. First, we show that the performance of previous approaches that combine Markov Random Fields with neural networks can be modestly improved by leveraging more powerful methods such as iterative join graph propagation, integer linear programming, and  regularization-based structure learning. Then we propose a new modeling framework called deep dependency networks, which augments a dependency network, a model that is easy to train and learns more accurate dependencies but is limited to Gibbs sampling for inference, to the output layer of a neural network. We show that despite its simplicity, jointly learning this new architecture yields significant improvements in performance over the baseline neural network. In particular, our experimental evaluation on three video activity classification datasets: Charades, Textually Annotated Cooking Scenes (TACoS), and Wetlab, and three multi-label image classification datasets: MS-COCO, PASCAL VOC, and NUS-WIDE show that deep dependency networks are almost always superior to pure neural architectures that do not use dependency networks.
  18. Put on your detective hat: What’s wrong in this video?
    Peddi, Rohith and Arya, Shivvrat and Challa, Bharath and Pallapothula, Likhitha and Vyas, Akshay and Zhang, Qifan and Wang, Jikai and Komaragiri, Vasundhara and Ruozzi, Nicholas and Ragan, Eric and Xiang, Yu and Gogate, Vibhav
    DMLR Data-centric Machine Learning Research Workshop, 2023
    PDF
  19. Chauhan, Vikas and Tiwari, Aruna and Arya, Shivvrat
    2020 International Joint Conference on Neural Networks (IJCNN), 2020
    DOI
    Abstract
    In this paper, a kernelized version of the random vector functional link network is proposed for multi-label classification. This classifier uses pseudo-inverse to find output weights of the network. As pseudo-inverse is non-iterative in nature, it requires less fine-tuning to train the network. Kernelization of RVFL makes it robust and stable as no need to tune the number of neuron in the enhancement layer. A threshold function is used with a kernelized random vector functional link network to make it suitable for multi-label learning problems. Experiments performed on three benchmark multi-label datasets bibtex, emotions, and scene shows that proposed classifier outperforms various the existing multi-label classifiers.