arxiv.org/abs/2602.22911v2
Low-Rank Adaptation (LoRA) dominates parameter-efficient fine-tuning (PEFT). However, it faces a critical ``linear ceiling'' in complex reasoning tasks: simply increasing the rank yields diminishing returns due to intrinsic linear constraints. We int...
arxiv.org/abs/2410.09771v2
Efficient neural networks are essential for scaling machine learning models to real-time applications and resource-constrained environments. Fully-connected feedforward layers (FFLs) introduce computation and parameter count bottlenecks within neural...
arxiv.org/abs/2601.22563v2
Efficient neural networks are essential for scaling machine learning models to real-time applications and resource-constrained environments. Fully-connected feedforward layers (FFLs) introduce computation and parameter count bottlenecks within neural...
www.reddit.com/r/LocalLLaMA/comments/1nqfck4/how_accurate_is_the_mteb_leaderboard/
It's weird how some 600m-1b parameter embedding beat other models like voyage-3-lg. Also how it doesn't even mention models like voyage-context-3....
arxiv.org/abs/0806.3050v2
The impact of a drop onto a liquid layer and the subsequent splash has important implications for diverse physical processes such as air-sea gas transfer, cooling, and combustion. In the {\it crown splash} parameter regime, the splash pattern is hi...
arxiv.org/abs/cond-mat/9504056v1
The static free energy of glassy systems can be expressed in terms of the Parisi order parameter function. When this function has a discontinuity, the location of the step is determined by maximizing the free energy. In dynamics a transition is fou...
arxiv.org/abs/1511.06774v1
Graph burning is a model for the spread of social contagion. The burning number is a graph parameter associated with graph burning that measures the speed of the spread of contagion in a graph; the lower the burning number, the faster the contagion s...
arxiv.org/abs/1706.03106v2
In this paper we study the graph parameter of burning number, introduced by Bonato, Janssen, and Roshanbin (2014). We are particular interested in determining the burning number of Circulant graphs. In this paper, we find upper and lower bounds on th...
arxiv.org/abs/2009.10642v1
Graph burning is a deterministic, discrete-time process that models how influence or contagion spreads in a graph. Associated to each graph is its burning number, which is a parameter that quantifies how quickly the influence spreads. We survey resul...
arxiv.org/abs/1906.11700v1
Probabilistic programming languages and other machine learning applications often require samples to be generated from a categorical distribution where the probability of each one of $n$ categories is specified as a parameter. If the parameters are h...
arxiv.org/abs/1111.6500v1
We study the binding energy, root-mean-square radius and quadrupole deformation parameter for the synthesized superheavy element Z = 115, within the formalism of relativistic mean field theory. The calculation is dones for various isotopes of Z = 115...
arxiv.org/abs/1111.5097v1
We study the solvability of a system of ordinary differential equations derived from null geodesics of the LTB metric with data given in terms of a so-called redshift parameter. Data is introduced along these geodesics by the luminosity distance func...
www.bing.com/ck/a?!&&p=6450a53ec7d94dbf6ff1602241a08d3e502e449c52742e4a2c06ae61587cb35cJmltdHM9MTc3Mjg0MTYwMA&ptn=3&ver=2&hsh=4&fclid=3002bbb2-888d-6c5e-3528-aca789306dc6&u=a1aHR0cHM6Ly9haS5zdGFja2V4Y2hhbmdlLmNvbS9xdWVzdGlvbnMvNTc2OS9pbi1hLWNubi1kb2VzLWVhY2gtbmV3LWZpbHRlci1oYXZlLWRpZmZlcmVudC13ZWlnaHRzLWZvci1lYWNoLWlucHV0LWNoYW5uZWwtb3I&ntb=1
Typically for a CNN architecture, in a single filter as described by your number_of_filters parameter, there is one 2D kernel per input channel. There are input_channels * number_of_filters sets of …
www.bing.com/ck/a?!&&p=54ac77d1813e19fa359a0fc2101a91a7f63a96e8e81a4e659bdac07ab39b61b2JmltdHM9MTc3Mjg0MTYwMA&ptn=3&ver=2&hsh=4&fclid=1f21d793-f1cd-63eb-0981-c086f0836203&u=a1aHR0cHM6Ly9raW1pLm1vb25zaG90LmNuL3NoYXJlL2N1c3VnYmVuM21rMmE2ODM3cDcw&ntb=1
Try Kimi K2 now, the open‑source trillion‑parameter MoE model, smarter coding, and agentic task automation.
en.wikipedia.org/wiki/Strain_hardening_exponent
The strain hardening exponent (also called the strain hardening index), usually denoted n {\displaystyle n} , is a measured parameter that quantifies
arxiv.org/abs/2204.06745v1
We introduce GPT-NeoX-20B, a 20 billion parameter autoregressive language model trained on the Pile, whose weights will be made freely and openly available to the public through a permissive license. It is, to the best of our knowledge, the largest d...
arxiv.org/abs/1702.07481v2
The Cooperative Patent Classifications (CPC) jointly developed by the European and US Patent Offices provide a new basis for mapping and portfolio analysis. This update provides an occasion for rethinking the parameter choices. The new maps are signi...
arxiv.org/abs/1309.6086v2
The parameter space of the simplest extension of the standard model, is studied in the light of the 125 GeV Higgs boson discovery. The Hill model extends the scalar sector of the standard model with a real singlet, that mixes with the SM Higgs boson....
arxiv.org/abs/1405.6619v1
According to the method of series rearrangement, we establish two generalizations of Andrews' curious $q$-series identity with an extra integer parameter. The limiting cases of them produce two extensions of Andrews' curious $_3F_2(\frac{3}{4})$-seri...
www.bing.com/ck/a?!&&p=815767d632ef9d71f29f676393563cecd11346790b516913d048dc89a4e3a1e4JmltdHM9MTc3Mjg0MTYwMA&ptn=3&ver=2&hsh=4&fclid=0dbb17ac-f0dc-6a15-07ae-00b9f1e96b4d&u=a1aHR0cHM6Ly9haS5zdGFja2V4Y2hhbmdlLmNvbS9xdWVzdGlvbnMvNTc2OS9pbi1hLWNubi1kb2VzLWVhY2gtbmV3LWZpbHRlci1oYXZlLWRpZmZlcmVudC13ZWlnaHRzLWZvci1lYWNoLWlucHV0LWNoYW5uZWwtb3I&ntb=1
Typically for a CNN architecture, in a single filter as described by your number_of_filters parameter, there is one 2D kernel per input channel. There are input_channels * number_of_filters sets of …