共找到 20 条结果
In the digital age, e-commerce has transformed the way consumers shop, offering convenience and accessibility. Nevertheless, concerns about the privacy and security of personal information shared on these platforms have risen. In this work, we investigate user privacy violations, noting the risks of data leakage to third-party entities. Utilizing a semi-automated data collection approach, we examine a selection of popular online e-shops, revealing that nearly 30% of them violate user privacy by disclosing personal information to third parties. We unveil how minimal user interaction across multiple e-commerce websites can result in a comprehensive privacy breach. We observe significant data-sharing patterns with platforms like Facebook, which use personal information to build user profiles and link them to social media accounts.
NASA engineers have found a clever way to squeeze more life out of Voyager 2, nearly half a century after it left Earth。 In an operation nicknamed the “Big Bang,” the team simultaneously shut down certain power-hungry hardware and switched to lower-power alternatives while keeping the spacecraft warm enough to survive
Classical approximation theorems ask for a new neural network whenever the target accuracy is improved. This paper studies the opposite possibility: can the network be chosen once and for all, and can accuracy be bought only by letting it run longer? We prove that this is possible for every continuous function on [-1,1]. More precisely, each such function is uniformly approximated by the time evolution of a single ReLU recurrent neural network with fixed weights and fixed hidden dimension. The mechanism behind the construction is a new intermediate model, the Turing machine with neural units (TMNU). This model retains the algorithmic freedom needed to implement polynomial approximation schemes, while remaining rigid enough to be simulated by RNNs with explicit bounds on hidden dimension and weight magnitude. The resulting convergence rates reflect the underlying polynomial approximation rates. We complement the construction with minimax lower bounds showing that runtime is not merely a proof artifact, but an unavoidable resource in this fixed-network approximation paradigm.
Conflicting objectives are general in RL alignment, and training on them data-efficiently is hard. Training a safety guard with RL means optimizing two objectives that conflict: catch real harm, and do not refuse benign prompts. Our finding is that over-refusal improves 22.4% to 12.8%, while under-refusal on adversarial attacks silently worsens 0.27 to 0.33. We present C-Guard, a constitution-grid instrument that generates the RL training data, and C-LIM, a per-cell learnability score that decides each cell's move: prune, densify, amend, expand. C-LIM flags the dead-weight data region before any training budget is spent: 187 untargeted rows had bought zero gain, and our method lifts the same region's learning impact 0.733 to 0.80. Code and the constitution are open-sourced.
Large vision-language models (LVLMs) hallucinate: they assert visual details that the image does not support. A principled remedy is selective prediction with a distribution-free guarantee-verify each claim and abstain when the claim is not grounded, so that the hallucination rate among asserted claims is provably bounded. We show, however, that this guarantee is bought at a brutal price: to keep the hallucination rate below $5\%$ on a balanced object-existence benchmark, a state-of-the-art conformal filter must abstain on more than $80\%$ of claims. We argue that abstention is wasteful when more visual evidence is cheaply available, and introduce Budgeted Conformal Evidence Acquisition (BCEA), which replaces the binary answer/abstain decision with a three-way choice: answer, abstain, or acquire additional visual evidence by re-examining the image (zooming, cropping, or applying a claim-specific intervention) under a bounded
Motivated by the prevalence of prediction problems in the economy, we study markets in which firms sell models to a consumer to help improve their prediction. Firms decide whether to enter, choose models to train on their data, and set prices. The consumer can purchase multiple models and use a weighted average of the models bought. Market outcomes can be expressed in terms of the \emph{bias-variance decompositions} of the models that firms sell. We give conditions when symmetric firms will choose different modeling techniques, e.g., each using only a subset of available covariates. We also show firms can choose inefficiently biased models or inefficiently costly models to deter entry by competitors.
We consider a discrete-time model of a financial market where a risky asset is bought and sold with transactions having a transient price impact. It is shown that the corresponding utility maximization problem admits a solution. We manage to remove some unnatural restrictions on the market depth and resilience processes that were present in earlier work. A non-standard feature of the problem is that the set of attainable portfolio values may fail the convexity property.
We consider a diffusion risk model where proportional reinsurance can be bought. In order to stabilise the surplus process, one tries to keep the drawdown, that is the difference of the surplus to its historical maximum, in an interval $[0,d)$. The observation times of the drawdowns form a renewal process. The retention levels can only be changed at the observation times either. We show that an optimal strategy exists and how it is determined. We illustrate the findings in the case of Poissonian observation times and deterministic inter-observation times.
The number of air transportation passengers during the holidays in Brazil has grown notably since the late nineties. One of the reasons is greater competition in airfares made possible by economic liberalization. This paper presents an econometric model of airline pricing aiming at estimating the impacts of holiday periods on fares, with special emphasis on three-day holiday events. It makes use of a database with daily collected data from the internet between 2008 and 2010 for the major Brazilian city, Sao Paulo. The econometric panel data model employs a two-way error components "within" estimator, controlling for airline/airport-pair fixed effect along with quotation and departure months effects. The decomposition of time effects between quotation and departure month effects is the main methodological contribution of the paper. Results allow for a comparative analysis of the performance of Sao Paulo's downtown and international airports - respectively, Congonhas (CGH), and Guarulhos (GRU) airports. As a result, the price of tickets bought 60 days in advance for flights with two stops leaving from the downtown airport fell by most.
Citations are widely considered in scientists' evaluation. As such, scientists may be incentivized to inflate their citation counts. While previous literature has examined self-citations and citation cartels, it remains unclear whether scientists can purchase citations. Here, we compile a dataset of ~1.6 million profiles on Google Scholar to examine instances of citation fraud on the platform. We survey faculty at highly-ranked universities, and confirm that Google Scholar is widely used when evaluating scientists. Intrigued by a citation-boosting service that we unravelled during our investigation, we contacted the service while undercover as a fictional author, and managed to purchase 50 citations. These findings provide conclusive evidence that citations can be bought in bulk, and highlight the need to look beyond citation counts.
Amazon is the world number one online retailer and has nearly every product a person could need along with a treasure trove of product reviews to help consumers make educated purchases. Companies want to find a way to increase their sales in a very crowded market, and using this data is key. A very good indicator of how a product is selling is its sales rank; which is calculated based on all-time sales of a product where recent sales are weighted more than older sales. Using the data from the Amazon products and reviews we determined that the most influential factors in determining the sales rank of a product were the number of products Amazon showed that other customers also bought, the number of products Amazon showed that customers also viewed, and the price of the product. These results were consistent for the Digital Music category, the Office Products category, and the subcategory Holsters under Cell Phones and Accessories.
In this paper, we consider a mean field model of social behavior where there are an infinite number of players, each of whom observes a type privately that represents her preference, and publicly observes a mean field state of types and actions of the players in the society. The types (and equivalently preferences) of the players are dynamically evolving. Each player is fully rational and forward-looking and makes a decision in each round t to buy a product. She receives a higher utility if the product she bought is aligned with her current preference and if there is a higher fraction of people who bought that product (thus a game of strategic complementarity). We show that for certain parameters when the weight of strategic complementarity is high, players eventually herd towards one of the actions with probability 1 which is when each player buys a product irrespective of her preference.
Recommender systems influence many of our interactions in the digital world -- impacting how we shop for clothes, sorting what we see when browsing YouTube or TikTok, and determining which restaurants and hotels we are shown when using hospitality platforms. Modern recommender systems are large, opaque models trained on a mixture of proprietary and open-source datasets. Naturally, issues of trust arise on both the developer and user side: is the system working correctly, and why did a user receive (or not receive) a particular recommendation? Providing an explanation alongside a recommendation alleviates some of these concerns. The status quo for auxiliary recommender system feedback is either user-specific explanations (e.g., "users who bought item B also bought item A") or item-specific explanations (e.g., "we are recommending item A because you watched/bought item B"). However, users bring personalized context into their search experience, valuing an item as a function of that item's attributes and their own personal preferences. In this work, we propose RecXplainer, a novel method for generating fine-grained explanations based on a user's preferences over the attributes of reco
Word co-occurrence patterns in language corpora contain a surprising amount of conceptual knowledge. Large language models (LLMs), trained to predict words in context, leverage these patterns to achieve impressive performance on diverse semantic tasks requiring world knowledge. An important but understudied question about LLMs' semantic abilities is whether they acquire generalized knowledge of common events. Here, we test whether five pre-trained LLMs (from 2018's BERT to 2023's MPT) assign higher likelihood to plausible descriptions of agent-patient interactions than to minimally different implausible versions of the same event. Using three curated sets of minimal sentence pairs (total n=1,215), we found that pre-trained LLMs possess substantial event knowledge, outperforming other distributional language models. In particular, they almost always assign higher likelihood to possible vs. impossible events (The teacher bought the laptop vs. The laptop bought the teacher). However, LLMs show less consistent preferences for likely vs. unlikely events (The nanny tutored the boy vs. The boy tutored the nanny). In follow-up analyses, we show that (i) LLM scores are driven by both plausi
The text generated by large language models is commonly controlled by prompting, where a prompt prepended to a user's query guides the model's output. The prompts used by companies to guide their models are often treated as secrets, to be hidden from the user making the query. They have even been treated as commodities to be bought and sold on marketplaces. However, anecdotal reports have shown adversarial users employing prompt extraction attacks to recover these prompts. In this paper, we present a framework for systematically measuring the effectiveness of these attacks. In experiments with 3 different sources of prompts and 11 underlying large language models, we find that simple text-based attacks can in fact reveal prompts with high probability. Our framework determines with high precision whether an extracted prompt is the actual secret prompt, rather than a model hallucination. Prompt extraction from real systems such as Claude 3 and ChatGPT further suggest that system prompts can be revealed by an adversary despite existing defenses in place.
Non-fungible tokens or NFTs are the digital assets on a blockchain. NFTs are unique and they cannot be divided like cryptocurrencies. NFTs could store digital ownership of an artwork or collections or can be fan tokens or tickets for clubs. NFTs are based on a smart contract on a blockchain network which supports them, such as Ethereum, Cardano or Polkadot. Most of the NFTs are now minted on Ethereum (ERC-20) network, but it has some main issues like high transaction fees and low speed. There are lots of domains which can be benefited from NFT technology such as art, music, gaming, sport and wildlife conservation. NFTs could be also bought or sold on lots of NFT marketplaces such as OpenSea and Chiliz. The trend is in a huge hype because the market cap and popularity of NFTs are growing significantly.
In this work, we propose a simple stochastic agent-based model to describe the revenue dynamics of a nightclub venue based on the relationship between profit and spatial occupation. The system consists of an underlying square lattice (nightclub's dance floor) where every attendee (agent) is allowed to move to its first neighboring cells. Each guess has a characteristic delayed time between drinks, denoted as $τ$, after which it will show an urge to drink. At this moment, the attendee will tend to move towards the bar where a drink will be bought. After it has left the bar zone, $τ$ time steps should pass so it shows once again the need to drink. Our model among other points show that it is no use filling the bar to obtain profit, and optimization should be analyzed. This can be done in a more secure way taking into consideration the ratio between income and ticket cost.
The goal of the classic football-pool problem is to determine how many lottery tickets are to be bought in order to guarantee at least $n-r$ correct guesses out of a sequence of $n$ games played. We study a generalized (second-order) version of this problem, in which any of these $n$ games consists of two sub-games. The second-order version of the football-pool problem is formulated using the notion of generalized-covering radius, recently proposed as a fundamental property of linear codes. We consider an extension of this property to general (not necessarily linear) codes, and provide an asymptotic solution to our problem by finding the optimal rate function of second-order covering codes given a fixed normalized covering radius. We also prove that the fraction of second-order covering codes among codes of sufficiently large rate tends to $1$ as the code length tends to $\infty$.
The problem of detecting terms that can be interesting to the advertiser is considered. If a company has already bought some advertising terms which describe certain services, it is reasonable to find out the terms bought by competing companies. A part of them can be recommended as future advertising terms to the company. The goal of this work is to propose better interpretable recommendations based on FCA and association rules.
In this study, regional (cities, towns and villages) data and tweet data are obtained from Twitter, and extract information of purchase information (Where and what bought) from the tweet data by morphological analysis and rule-based dependency analysis. Then, the "The regional information" and "The information of purchase history (Where and what bought information)" are captured as bipartite graph, and Responsiveness Pair Clustering analysis (a clustering using correspondence analysis as similarity measure) is conducted. In this study, since it was found to be difficult to analyze a network such as bipartite graph having limitations in links by using modularity Q, responsiveness is used instead of modularity Q as similarity measure. As a result of this analysis, "regional information cluster" which refers to similar "The information of purchase history" nodes group is generated. Finally, similar regions are visualized by mapping the regional information cluster on the map. This visualization system is expected to contribute as an analytical tool for customers purchasing behavior and so on.