We present DuetGen, a novel framework for generating interactive two-person dances from music. The key challenge of this task lies in the inherent complexities of two-person dance interactions, where the partners need to synchronize both with each other and with the music. Inspired by the recent advances in motion synthesis, we propose a two-stage solution: encoding two-person motions into discrete tokens and then generating these tokens from music. To effectively capture intricate interactions, we represent both dancers' motions as a unified whole to learn the necessary motion tokens, and adopt a coarse-to-fine learning strategy in both the stages. Our first stage utilizes a VQ-VAE that hierarchically separates high-level semantic features at a coarse temporal resolution from low-level details at a finer resolution, producing two discrete token sequences at different abstraction levels. Subsequently, in the second stage, two generative masked transformers learn to map music signals to these dance tokens: the first producing high-level semantic tokens, and the second, conditioned on music and these semantic tokens, producing the low-level tokens. We train both transformers to learn
Modeling human-human interactions from text remains challenging because it requires not only realistic individual dynamics but also precise, text-consistent spatiotemporal coupling between agents. Currently, progress is hindered by 1) limited two-person training data, inadequate to capture the diverse intricacies of two-person interactions; and 2) insufficiently fine-grained text-to-interaction modeling, where language conditioning collapses rich, structured prompts into a single sentence embedding. To address these limitations, we propose our Text2Interact framework, designed to generate realistic, text-aligned human-human interactions through a scalable high-fidelity interaction data synthesizer and an effective spatiotemporal coordination pipeline. First, we present InterCompose, a scalable synthesis-by-composition pipeline that aligns LLM-generated interaction descriptions with strong single-person motion priors. Given a prompt and a motion for an agent, InterCompose retrieves candidate single-person motions, trains a conditional reaction generator for another agent, and uses a neural motion evaluator to filter weak or misaligned samples-expanding interaction coverage without e
We consider a sub-class of bi-matrix games which we refer to as two-person (hereafter referred to as two-player) additively-separable sum (TPASS) games, where the sum of the pay-offs of the two players is additively separable. The row player's pay-off at each pair of pure strategies, is the sum of two numbers, the first of which may be dependent on the pure strategy chosen by the column player and the second being independent of the pure strategy chosen by the column player. The column player's pay-off at each pair of pure strategies, is also the sum of two numbers, the first of which may be dependent on the pure strategy chosen by the row player and the second being independent of the pure strategy chosen by the row player. The sum of the inter-dependent components of the pay-offs of the two players is assumed to be zero. We prove the existence of equilibrium for such games and show that the set of equilibria for such games is the projection on the set of strategy pairs of the solutions of a pair of linear programming problems that are dual to each other. This result is a generalization of the corresponding and well-known result for two-person zero-sum games. We also show that a (
Generating realistic human motion with high-level controls is a crucial task for social understanding, robotics, and animation. With high-quality MOCAP data becoming more available recently, a wide range of data-driven approaches have been presented. However, modelling multi-person interactions still remains a less explored area. In this paper, we present Graph-driven Interaction Sampling, a method that can generate realistic and diverse multi-person interactions by leveraging existing two-person motion diffusion models as motion priors. Instead of training a new model specific to multi-person interaction synthesis, our key insight is to spatially and temporally separate complex multi-person interactions into a graph structure of two-person interactions, which we name the Pairwise Interaction Graph. We thus decompose the generation task into simultaneous single-person motion generation conditioned on one other's motion. In addition, to reduce artifacts such as interpenetrations of body parts in generated multi-person interactions, we introduce two graph-dependent guidance terms into the diffusion sampling scheme. Unlike previous work, our method can produce various high-quality mul
A player's payoff is modeled as consisting of two parts: a rational-value part and a distortion-value part. It is argued that the (total) payoff function should be used to explain and predict the behaviors of the players, while the rational value function should be used to conduct welfare analysis of the final outcome. We use the Nash demand game to illustrate our model.
As a fundamental aspect of human life, two-person interactions contain meaningful information about people's activities, relationships, and social settings. Human action recognition serves as the foundation for many smart applications, with a strong focus on personal privacy. However, recognizing two-person interactions poses more challenges due to increased body occlusion and overlap compared to single-person actions. In this paper, we propose a point cloud-based network named Two-stream Multi-level Dynamic Point Transformer for two-person interaction recognition. Our model addresses the challenge of recognizing two-person interactions by incorporating local-region spatial information, appearance information, and motion information. To achieve this, we introduce a designed frame selection method named Interval Frame Sampling (IFS), which efficiently samples frames from videos, capturing more discriminative information in a relatively short processing time. Subsequently, a frame features learning module and a two-stream multi-level feature aggregation module extract global and partial features from the sampled frames, effectively representing the local-region spatial information, a
Conversational scenarios are very common in real-world settings, yet existing co-speech motion synthesis approaches often fall short in these contexts, where one person's audio and gestures will influence the other's responses. Additionally, most existing methods rely on offline sequence-to-sequence frameworks, which are unsuitable for online applications. In this work, we introduce an audio-driven, auto-regressive system designed to synthesize dynamic movements for two characters during a conversation. At the core of our approach is a diffusion-based full-body motion synthesis model, which is conditioned on the past states of both characters, speech audio, and a task-oriented motion trajectory input, allowing for flexible spatial control. To enhance the model's ability to learn diverse interactions, we have enriched existing two-person conversational motion datasets with more dynamic and interactive motions. We evaluate our system through multiple experiments to show it outperforms across a variety of tasks, including single and two-person co-speech motion generation, as well as interactive motion generation. To the best of our knowledge, this is the first system capable of genera
The observation that every two-person adversarial game is an affine transformation of a zero-sum game is traceable to Luce & Raiffa (1957) and made explicit in Aumann (1987). Recent work of (ADP) Adler et al. (2009), and of Raimondo (2023) in increasing generality, proves what has so far remained a conjecture. We present two proofs of an even more general formulation: the first draws on multilinear utility theory developed by Fishburn & Roberts (1978); the second is a consequence of the ADP proof itself for a special case of a two-player game with a set of three actions.
This paper addresses a class of two-person zero-sum stochastic differential equations, which encompass Markov chains and fractional Brownian motion, and satisfy some monotonicity conditions over an infinite time horizon. Within the framework of forward-backward stochastic differential equations (FBSDEs) that describe system evolution, we extend the classical It$\rm\hat{o}$'s formula to accommodate complex scenarios involving Brownian motion, fractional Brownian motion, and Markov chains simultaneously. By applying the Banach fixed-point theorem and approximation methods respectively, we theoretically guarantee the existence and uniqueness of solutions for FBSDEs in infinite horizon. Furthermore, we apply the method for the first time to the optimal control problem in a two-player zero-sum game, deriving the optimal control strategies for both players by solving the FBSDEs system. Finally, we conduct an analysis of the impact of the cross-term $S(\cdot)$ in the cost function on the solution, revealing its crucial role in the optimization process.
This paper presents a pioneering investigation into discrete-time two-person non-zero-sum linear quadratic (LQ) stochastic games with random coefficients. We derive necessary and sufficient conditions for the existence of open-loop Nash equilibria using convex variational calculus. To obtain explicit expressions for the Nash equilibria, we introduce fully coupled forward-backward stochastic difference equations (FBS$Δ$E, for short), which provide a dual characterization of these Nash equilibria. Additionally, we develop non-symmetric stochastic Riccati equations that decouple the stochastic Hamiltonian system for each player, enabling the derivation of closed-loop feedback forms for open-loop Nash equilibrium strategies. A notable aspect of this research is the complete randomness of the coefficients, which results in the corresponding Riccati equations becoming fully nonlinear higher-order backward stochastic difference equations. It distinguishes our non-zero-sum difference game from the deterministic case, where the Riccati equations reduce to algebraic forms.
We prove that every finite two-person shortest path game, where the local cost of every move is positive for each player, has a Nash equilibrium (NE) in pure stationary strategies, which can be computed in polynomial time. We also extend the existence result to infinite graphs with finite out-degrees. Moreover, our proof gives that a terminal NE (in which the play is a path from the initial position to a terminal) exists provided at least one of the two players can guarantee reaching a terminal. If none of the players can do it, in other words, if each of the two players has a strategy that separates all terminals from the initial position $s$, then, obviously, a cyclic NE exists, although its cost is infinite for both players, since we restrict ourselves to positive games. We conjecture that a terminal NE exists too, provided there exists a directed path from $s$ to a terminal. However, this is open. We extend our result to short paths interdiction games, where at each vertex, we allow one player to block some of the arcs and the other player to choose one of the non-blocked arcs. Assuming that blocking sets are chosen from an independence system given by an oracle, we give an alg
Graph convolutional networks (GCNs) have been the predominant methods in skeleton-based human action recognition, including human-human interaction recognition. However, when dealing with interaction sequences, current GCN-based methods simply split the two-person skeleton into two discrete graphs and perform graph convolution separately as done for single-person action classification. Such operations ignore rich interactive information and hinder effective spatial inter-body relationship modeling. To overcome the above shortcoming, we introduce a novel unified two-person graph to represent inter-body and intra-body correlations between joints. Experiments show accuracy improvements in recognizing both interactions and individual actions when utilizing the proposed two-person graph topology. In addition, We design several graph labeling strategies to supervise the model to learn discriminant spatial-temporal interactive features. Finally, we propose a two-person graph convolutional network (2P-GCN). Our model achieves state-of-the-art results on four benchmarks of three interaction datasets: SBU, interaction subsets of NTU-RGB+D and NTU-RGB+D 120.
This paper presents a time-invariant network flow model capturing two-person ride-pooling that can be integrated within design and planning frameworks for Mobility-on-Demand systems. In these type of models, the arrival process of travel requests is described by a Poisson process, meaning that there is only statistical insight into request times, including the probability that two requests may be pooled together. Taking advantage of this feature, we devise a method to capture ride-pooling from a stochastic mesoscopic perspective. This way, we are able to transform the original set of requests into an equivalent set including pooled ones which can be integrated within standard network flow problems, which in turn can be efficiently solved with off-the-shelf LP solvers for a given ride-pooling request assignment. Thereby, to compute such an assignment, we devise a polynomial-time algorithm that is optimal w.r.t. an approximated version of the problem. Finally, we perform a case study of Sioux Falls, South Dakota, USA, where we quantify the effects that waiting time and experienced delay have on the vehicle-hours traveled. Our results suggest that the higher the demands per unit time,
We explore a version of the minimax theorem for two-person win-lose games with infinitely many pure strategies. In the countable case, we give a combinatorial condition on the game which implies the minimax property. In the general case, we prove that a game satisfies the minimax property along with all its subgames if and only if none of its subgames is isomorphic to the "larger number game." This generalizes a recent theorem of Hanneke, Livni and Moran. We also propose several applications of our results outside of game theory.
Current approaches for 3D human motion synthesis generate high quality animations of digital humans performing a wide variety of actions and gestures. However, a notable technological gap exists in addressing the complex dynamics of multi human interactions within this paradigm. In this work, we present ReMoS, a denoising diffusion based model that synthesizes full body reactive motion of a person in a two person interaction scenario. Given the motion of one person, we employ a combined spatio temporal cross attention mechanism to synthesize the reactive body and hand motion of the second person, thereby completing the interactions between the two. We demonstrate ReMoS across challenging two person scenarios such as pair dancing, Ninjutsu, kickboxing, and acrobatics, where one persons movements have complex and diverse influences on the other. We also contribute the ReMoCap dataset for two person interactions containing full body and finger motions. We evaluate ReMoS through multiple quantitative metrics, qualitative visualizations, and a user study, and also indicate usability in interactive motion editing applications.
Skeleton-based two-person interaction recognition has been gaining increasing attention as advancements are made in pose estimation and graph convolutional networks. Although the accuracy has been gradually improving, the increasing computational complexity makes it more impractical for a real-world environment. There is still room for accuracy improvement as the conventional methods do not fully represent the relationship between inter-body joints. In this paper, we propose a lightweight model for accurately recognizing two-person interactions. In addition to the architecture, which incorporates middle fusion, we introduce a factorized convolution technique to reduce the weight parameters of the model. We also introduce a network stream that accounts for relative distance changes between inter-body joints to improve accuracy. Experiments using two large-scale datasets, NTU RGB+D 60 and 120, show that our method simultaneously achieved the highest accuracy and relatively low computational complexity compared with the conventional methods.
We consider two-person bargaining problems in which (only) the disagreement outcome is private (and possibly correlated) information and it is common knowledge that disagreement is inefficient. We show that if the Pareto frontier is linear, the outcome of an ex post efficient mechanism cannot depend on the disagreement payoffs. If the frontier is non-linear, the result continues to hold when the disagreement payoffs are independent or there is a player with at most two types. We discuss implications of these results for axiomatic bargaining theory and for full surplus extraction in mechanism design.
We consider finite two-person normal form games. The following four properties of their game forms are equivalent: (i) Nash-solvability, (ii) zero-sum-solvability, (iii) win-lose-solvability, and (iv) tightness. For (ii, iii, iv) this was shown by Edmonds and Fulkerson in 1970. Then, in 1975, (i) was added to this list and it was also shown that these results cannot be generalized for $n$-person case with $n > 2$. In 1990, tightness was extended to vector game forms ($v$-forms) and it was shown that such $v$-tightness and zero-sum-solvability are still equivalent, yet, do not imply Nash-solvability. These results are applicable to several classes of stochastic games with perfect information. Here we suggest one more extension of tightness introducing $v^+$-tight vector game forms ($v^+$-forms). We show that such $v^+$-tightness and Nash-solvability are equivalent in case of weakly rectangular game forms and positive cost functions. This result allows us to reduce the so-called bi-shortest path conjecture to $v^+$-tightness of $v^+$-forms. However, both (equivalent) statements remain open.
Amid the surge in generic text-to-video generation, the field of personalized human video generation has witnessed notable advancements, primarily concentrated on single-person scenarios. However, to our knowledge, the domain of two-person interactions, particularly in the context of martial arts combat, remains uncharted. We identify a significant gap: existing models for single-person dancing generation prove insufficient for capturing the subtleties and complexities of two engaged fighters, resulting in challenges such as identity confusion, anomalous limbs, and action mismatches. To address this, we introduce a pioneering new task, Personalized Martial Arts Combat Video Generation. Our approach, MagicFight, is specifically crafted to overcome these hurdles. Given this pioneering task, we face a lack of appropriate datasets. Thus, we generate a bespoke dataset using the game physics engine Unity, meticulously crafting a multitude of 3D characters, martial arts moves, and scenes designed to represent the diversity of combat. MagicFight refines and adapts existing models and strategies to generate high-fidelity two-person combat videos that maintain individual identities and ensur
Large Language Model (LLM)-based mobile agents have made significant performance advancements. However, these agents often follow explicit user instructions while overlooking personalized needs, leading to significant limitations for real users, particularly without personalized context: (1) inability to interpret ambiguous instructions, (2) lack of learning from user interaction history, and (3) failure to handle personalized instructions. To alleviate the above challenges, we propose Me-Agent, a learnable and memorable personalized mobile agent. Specifically, Me-Agent incorporates a two-level user habit learning approach. At the prompt level, we design a user preference learning strategy enhanced with a Personal Reward Model to improve personalization performance. At the memory level, we design a Hierarchical Preference Memory, which stores users' long-term memory and app-specific memory in different level memory. To validate the personalization capabilities of mobile agents, we introduce User FingerTip, a new benchmark featuring numerous ambiguous instructions for daily life. Extensive experiments on User FingerTip and general benchmarks demonstrate that Me-Agent achieves state-