共找到 20 条结果
A practical studio guide for creating 2K-class AI video with synced audio using MiniMax H3 (Hailuo 3.0 / 海螺3) text-to-video and image-to-video workflows. Independent third-party service at minimaxh3.art — not the official MiniMax/Hailuo site. Full guide: https://minimaxh3.art/blog/minimax-h3-2k-text-image-video-guide PDF: https://minimaxh3.art/blog/minimax-h3-2k-text-image-video-guide/minimax-h3-2k-text-image-video-guide.pdf Homepage: https://minimaxh3.art/
We study reinforcement learning with linear function approximation where the transition probability and reward functions are linear with respect to a feature mapping $\boldsymbolϕ(s,a)$. Specifically, we consider the episodic inhomogeneous linear Markov Decision Process (MDP), and propose a novel computation-efficient algorithm, LSVI-UCB$^+$, which achieves an $\widetilde{O}(Hd\sqrt{T})$ regret bound where $H$ is the episode length, $d$ is the feature dimension, and $T$ is the number of steps. LSVI-UCB$^+$ builds on weighted ridge regression and upper confidence value iteration with a Bernstein-type exploration bonus. Our statistical results are obtained with novel analytical tools, including a new Bernstein self-normalized bound with conservatism on elliptical potentials, and refined analysis of the correction term. This is a minimax optimal algorithm for linear MDPs up to logarithmic factors, which closes the $\sqrt{Hd}$ gap between the upper bound of $\widetilde{O}(\sqrt{H^3d^3T})$ in (Jin et al., 2020) and lower bound of $Ω(Hd\sqrt{T})$ for linear MDPs.
The matrix completion problem consists in reconstructing a matrix from a sample of entries, possibly observed with noise. A popular class of estimator, known as nuclear norm penalized estimators, are based on minimizing the sum of a data fitting term and a nuclear norm penalization. Here, we investigate the case where the noise distribution belongs to the exponential family and is sub-exponential. Our framework alllows for a general sampling scheme. We first consider an estimator defined as the minimizer of the sum of a log-likelihood term and a nuclear norm penalization and prove an upper bound on the Frobenius prediction risk. The rate obtained improves on previous works on matrix completion for exponential family. When the sampling distribution is known, we propose another estimator and prove an oracle inequality w.r.t. the Kullback-Leibler prediction risk, which translates immediatly into an upper bound on the Frobenius prediction risk. Finally, we show that all the rates obtained are minimax optimal up to a logarithmic factor.
We study reinforcement learning (RL) with linear function approximation where the underlying transition probability kernel of the Markov decision process (MDP) is a linear mixture model (Jia et al., 2020; Ayoub et al., 2020; Zhou et al., 2020) and the learning agent has access to either an integration or a sampling oracle of the individual basis kernels. We propose a new Bernstein-type concentration inequality for self-normalized martingales for linear bandit problems with bounded noise. Based on the new inequality, we propose a new, computationally efficient algorithm with linear function approximation named $\text{UCRL-VTR}^{+}$ for the aforementioned linear mixture MDPs in the episodic undiscounted setting. We show that $\text{UCRL-VTR}^{+}$ attains an $\tilde O(dH\sqrt{T})$ regret where $d$ is the dimension of feature mapping, $H$ is the length of the episode and $T$ is the number of interactions with the MDP. We also prove a matching lower bound $Ω(dH\sqrt{T})$ for this setting, which shows that $\text{UCRL-VTR}^{+}$ is minimax optimal up to logarithmic factors. In addition, we propose the $\text{UCLK}^{+}$ algorithm for the same family of MDPs under discounting and show that it attains an $\tilde O(d\sqrt{T}/(1-γ)^{1.5})$ regret, where $γ\in [0,1)$ is the discount factor. Our upper bound matches the lower bound $Ω(d\sqrt{T}/(1-γ)^{1.5})$ proved by Zhou et al. (2020) up to logarithmic factors, suggesting that $\text{UCLK}^{+}$ is nearly minimax optimal. To the best of our knowledge, these are the first computationally efficient, nearly minimax optimal algorithms for RL with linear function approximation.
We study various aspects related to boundary regularity of complete properly embedded Willmore surfaces in H3, particularly those related to assumptions on boundedness or smallness of a certain weighted version of the Willmore energy. We prove, in particular, that small energy controls C1 boundary regularity. We examine the possible lack of C1 convergence for sequences of surfaces with bounded Willmore energy and find that the mechanism responsible for this is a bubbling phenomenon, where energy escapes to infinity.
暂无摘要(点击查看原文获取完整内容)
暂无摘要(点击查看原文获取完整内容)
<h3>Background</h3> Direct immune stimulation using cytokines, including interleukin-2 (IL-2), drives immune-mediated cytotoxic cancer responses.<sup>1</sup> Toxicity of high-dose IL-2 limits its activity and clinical utility. Preclinical studies show that targeting IL-2 to specific cell subsets substantially increases its therapeutic index. In multiple cancers, tumor-infiltrating T-cells display higher levels of programmed cell death protein (PD-1) relative to circulating and tissue-resident T-cells,<sup>2–4</sup> making PD-1 a target for selective delivery of IL-2. TEV-56278 is a human antibody-cytokine fusion protein comprised of an antibody targeting PD-1 and attenuated IL-2. TEV-56278 is designed to deliver IL-2 selectively to PD-1+ T cells, thus amplifying anti-tumor T-cell activity while minimizing off-target systemic toxicities. Furthermore, its binding to PD-1 receptors does not block ligand binding and does not compete for binding with known PD1-blocking antibodies, allowing for use in combination. Here, we describe a phase 1a/1b trial design evaluating TEV-56278 as monotherapy and in combination with pembrolizumab. <h3>Methods</h3> This multicenter, open-label, Phase 1, first-in-human, dose-escalation, dose-expansion study is being conducted in 3 parts (figure 1). Patients >18 years of age with various locally advanced or metastatic solid tumors who have progressed on or were intolerant to standard of care therapies and resistant to anti-PD-(L)1 treatment will be included. The primary objective of dose escalation is to assess safety, tolerability and determine the recommended Phase 2 dose (RP2D) of TEV-56278 alone and in combination with pembrolizumab. Secondary objectives include assessment of pharmacokinetics (PK) and anti-tumor activity. Determination of the maximum tolerated dose will be based on a Bayesian optimal interval (BOIN) design.<sup>5</sup> Additional pharmacodynamic (PD) and biomarker data, including biomarkers of IL-2 signaling and activation of PD-1+ T cells, will be collected during backfill of monotherapy dose cohort(s). After determination of the monotherapy RP2D, 2 cohorts of patients with locally advanced or metastatic primary and secondary resistant melanoma (A&B) and 2 cohorts of primary and secondary resistant non-small cell lung cancer (NSCLC; [C&D]) will be accrued to evaluate the primary objective of anti-tumor activity. The Simon two-stage admissible design<sup>6–8</sup> will be used for Cohort B (RP2D and a dose <RP2D), while a Bayesian approach will be used for A, C, and D with RP2D. The objective response rate (ORR) and duration of response (DOR) will be assessed based on Response Evaluation Criteria in Solid Tumors (RECISTv1.1). Secondary objectives include safety, tolerability, PK and other efficacy measures. This study is open for participant enrollment, targeting up to 240 patients. <h3>Acknowledgements</h3> The study is funded by Teva Pharmaceuticals. Software that is freely available to the public was used for the BOIN design (mdanderson.org) and the Simon two-stage admissible design (unc.edu). <h3>Trial Registration</h3> ClinicalTrials. gov Identifier: NCT06480552. <h3>References</h3> Hanzly M, Aboumohamed A, Yarlagadda N, Creighton T, Digiorgio L, Fredrick A, <i>et al</i>. High-dose interleukin-2 therapy for metastatic renal cell carcinoma: a contemporary experience. <i>Urology</i> 2014;<b>83</b>(5):1129–34. Ahmadzadeh M, Johnson LA, Heemskerk B, Wunderlich JR, Dudley ME, White DE, <i>et al</i>. Tumor antigen-specific CD8 T-cells infiltrating the tumor express high levels of PD-1 and are functionally impaired. <i>Blood</i> 2009;<b>114</b>(8):1537–44. Badoual C, Hans S, Merillon N, Van Ryswick C, Ravel P, Benhamouda N, <i>et al</i>. PD-1-expressing tumor-infiltrating T-cells are a favorable prognostic biomarker in HPV-associated head and neck cancer. <i>Cancer Res</i> 2013;<b>73</b>(1):128–38. Gros A, Robbins PF, Yao X, Li YF, Turcotte S, Tran E,<i> et al</i>. PD-1 identifies the patient-specific CD8<sup>+</sup> tumor-reactive repertoire infiltrating human tumors. <i>J Clin Invest</i> 2014;<b>124</b>(5):2246–59. Liu S, Yuan Y. Bayesian optimal interval designs for phase I clinical trials. <i>J R, Stat Soc Ser C Appl Stat</i> 2015;<b>64</b>:507–23. Jung SH, Carey M, Kim KM. Graphical search for two-stage designs for phase ii clinical trials. <i>Control</i> 2001;<b>22</b>(4):367–72. Jung SH, Lee T, Kim KM, George SL. Admissible two-stage designs for phase II cancer clinical trials.<i> Stat Med</i> 2004;<b>23</b>(4):561–69. Qin F, Wu J, Chen F, Wei Y, Zhao Y, Jiang Z, <i>et al.</i> Optimal, minimax and admissible two-stage design for phase II oncology clinical trials. <i>BMC Med Res Methodol</i> 2020;<b>20</b>(1):126. <h3>Ethics Approval</h3> Central and local/site Institutional Review Board approval has been obtained prior to study initiation. All enrolled subjects must provide signed informed consent prior to any study related procedures being completed.
We study nonparametric methods for estimating the value function of an infinite-horizon discounted Markov reward process (MRP). We analyze the kernel-based least-squares temporal difference (LSTD) estimate, which can be understood either as a nonparametric instrumental variables method, or as a projected approximation to the Bellman fixed point equation. Our analysis imposes no assumptions on the transition operator of the Markov chain, but rather only conditions on the reward function and population-level kernel LSTD solutions. Using empirical process theory and concentration inequalities, we establish a nonasymptotic upper bound on the error with explicit dependence on the effective horizon H=(1−γ)−1 of the Markov reward process, the eigenvalues of the associated kernel operator, as well as the instance-dependent variance of the Bellman residual error. In addition, we prove minimax lower bounds over subclasses of MRPs, which shows that our guarantees are optimal in terms of the sample size n and the effective horizon H. Whereas existing worst-case theory predicts cubic scaling (H3) in the effective horizon, our theory reveals a much wider range of scalings, depending on the kernel, the stationary distribution, and the variance of the Bellman residual error. Notably, it is only parametric and near-parametric problems that can ever achieve the worst-case cubic scaling.
暂无摘要(点击查看原文获取完整内容)
We study the problem of high-dimensional Principal Component Analysis (PCA) with missing observations. In a simple, homogeneous observation model, we show that an existing observed-proportion weighted (OPW) estimator of the leading principal components can (nearly) attain the minimax optimal rate of convergence, which exhibits an interesting phase transition. However, deeper investigation reveals that, particularly in more realistic settings where the observation probabilities are heterogeneous, the empirical performance of the OPW estimator can be unsatisfactory; moreover, in the noiseless case, it fails to provide exact recovery of the principal components. Our main contribution, then, is to introduce a new method, which we call primePCA, that is designed to cope with situations where observations may be missing in a heterogeneous manner. Starting from the OPW estimator, primePCA iteratively projects the observed entries of the data matrix onto the column space of our current estimate to impute the missing entries, and then updates our estimate by computing the leading right singular space of the imputed data matrix. We prove that the error of primePCA converges to zero at a geometric rate in the noiseless case, and when the signal strength is not too small. An important feature of our theoretical guarantees is that they depend on average, as opposed to worst-case, properties of the missingness mechanism. Our numerical studies on both simulated and real data reveal that primePCA exhibits very encouraging performance across a wide range of scenarios, including settings where the data are not Missing Completely At Random.
Understanding statistical inference under possibly nonsparse high-dimensional models has gained much interest recently. For a given component of the regression coefficient, we show that the difficulty of the problem depends on the sparsity of the corresponding row of the precision matrix of the covariates, not the sparsity of the regression coefficients. We develop new concepts of uniform and essentially uniform nontestability that allow the study of limitations of tests across a broad set of alternatives. Uniform nontestability identifies a collection of alternatives such that the power of any test, against any alternative in the group, is asymptotically at most equal to the nominal size. Implications of the new constructions include new minimax testability results that, in sharp contrast to the current results, do not depend on the sparsity of the regression parameters. We identify new tradeoffs between testability and feature correlation. In particular, we show that, in models with weak feature correlations, minimax lower bound can be attained by a test whose power has the n rate, regardless of the size of the model sparsity.
The definitive version is available at www.blackwell-synergy.com Copyright Blackwell Publishing DOI : 10.1111/j.1365-2966.2007.11963.x
We present the induced generalized ordered weighted averaging (IGOWA) operator. It is a new aggregation operator that generalizes the OWA operator by using the main characteristics of two well known aggregation operators: the generalized OWA and the induced OWA operator. Then, this operator uses generalized means and order inducing variables in the reordering process. With this formulation, we get a wide range of aggregation operators that include all the particular cases of the IOWA and the GOWA operator, and a lot of other cases such as the induced ordered weighted geometric (IOWG) operator and the induced ordered weighted quadratic averaging (IOWQA) operator. We further generalize the IGOWA operator by using quasi-arithmetic means. The result is the Quasi-IOWA operator. Finally, we also develop a numerical example of the new approach in a financial decision making problem.
Bandwidth selection for procedures such as kernel density estimation and local regression have been widely studied over the past decade. Substantial “evidence” has been collected to establish superior performance of modern plug-in methods in comparison to methods such as cross validation; this has ranged from detailed analysis of rates of convergence, to simulations, to superior performance on real datasets. In this work we take a detailed look at some of this evidence, looking into the sources of differences. Our findings challenge the claimed superiority of plug-in methods on several fronts. First, plug-in methods are heavily dependent on arbitrary specification of pilot bandwidths and fail when this specification is wrong. Second, the often-quoted variability and undersmoothing of cross validation simply reflects the uncertainty of band-width selection; plug-in methods reflect this uncertainty by oversmoothing and missing important features when given difficult problems. Third, we look at asymptotic theory. Plug-in methods use available curvature information in an inefficient manner, resulting in inefficient estimates. Previous comparisons with classical approaches penalized the classical approaches for this inefficiency. Asymptotically, the plug-in based estimates are beaten by their own pilot estimates.
暂无摘要(点击查看原文获取完整内容)
Considered as rest points of ODE on L p , stationary viscous shock waves present a critical case for which standard semigroup methods do not suffice to determine stability. More precisely, there is no spectral gap between stationary modes and essential spectrum of the linearized operator about the wave, a fact that precludes the usual analysis by decomposition into invariant subspaces. For this reason, there have been until recently no results on shock stability from the semigroup perspective except in the scalar or totally compressive case ([Sat], [K.2], resp.), each of which can be reduced to the standard semigroup setting by Sattinger's method of weighted norms. We overcome this difficulty in the general case by the introduction of new, pointwise semigroup techniques, generalizing earlier work of Howard [H.1], Kapitula [K.1-2], and Zeng [Ze,LZe]. These techniques allow us to do "hard" analysis in PDE within the dynamical systems/semigroup framework: in particular, to obtain sharp, global pointwise bounds on the Green's function of the linearized operator around the wave, sufficient for the analysis of linear and nonlinear stability. The method is general, and should find applications also in other situations of sensitive stability.
Deep learning has recently achieved great success in many areas due to its strong capacity in data process. For instance, it has been widely used in financial areas such as stock market prediction, portfolio optimization, financial information processing and trade execution strategies. Stock market prediction is one of the most popular and valuable area in finance. In this paper, we propose a novel architecture of Generative Adversarial Network (GAN) with the Multi-Layer Perceptron (MLP) as the discriminator and the Long Short-Term Memory (LSTM) as the generator for forecasting the closing price of stocks. The generator is built by LSTM to mine the data distributions of stocks from given data in stock market and generate data in the same distributions, whereas the discriminator designed by MLP aims to discriminate the real stock data and generated data. We choose the daily data on S&P 500 Index and several stocks in a wide range of trading days and try to predict the daily closing price. Experimental results show that our novel GAN can get a promising performance in the closing price prediction on the real data compared with other models in machine learning and deep learning.
We consider functionals which are not bounded from above or from below even modulo compact perturbations, and which exhibit certain symmetries with respect to the action of a compact Lie group. We develop a method which permits us to prove the existence of multiple critical points for such functionals. The proofs are carried out directly in an infinite dimensional Hilbert space, and they are based on minimax arguments. The applications given here are to Hamiltonian systems of ordinary differential equations where the existence of multiple time-periodic solutions is established for several classes of Hamiltonians. Symmetry properties of these Hamiltonians such as time translation invariancy or evenness are exploited.
The multisource Weber problem is to locate simultaneously m facilities in the Euclidean plane to minimize the total transportation cost for satisfying the demand of n fixed users, each supplied from its closest facility. Many heuristics have been proposed for this problem, as well as a few exact algorithms. Heuristics are needed to solve quickly large problems and to provide good initial solutions for exact algorithms. We compare various heuristics, i.e., alternative location-allocation (Cooper 1964), projection (Bongartz et al. 1994), Tabu search (Brimberg and Mladenović 1996a), p-Median plus Weber (Hansen et al. 1996), Genetic search and several versions of Variable Neighbourhood search. Based on empirical tests that are reported, it is found that most traditional and some recent heuristics give poor results when the number of facilities to locate is large and that Variable Neighbourhood search gives consistently best results, on average, in moderate computing time.