arxiv.org/abs/2501.16211v1
Activities in underwater environments are paramount in several scenarios, which drives the continuous development of underwater image enhancement techniques. A major challenge in this domain is the depth at which images are captured, with increasing...
www.bing.com/ck/a?!&&p=6f36f71b889c6beccf1cdbeb1982630edc741e5c3ec58acd612bcafa3af30dddJmltdHM9MTc3MjY2ODgwMA&ptn=3&ver=2&hsh=4&fclid=1b1b2e57-fe98-64e3-3cdd-3943ffe565cc&u=a1aHR0cHM6Ly9vcGVuYWkuY29tL2luZGV4L3NvcmEv&ntb=1
Feb 15, 2024 · Sora is a diffusion model, which generates a video by starting off with one that looks like static noise and gradually transforms it by removing the noise over many steps.
www.bing.com/ck/a?!&&p=bb3c7d01c9df97724497fde0d1f3fab9f957326a9205bf530c33375420548123JmltdHM9MTc3MjY2ODgwMA&ptn=3&ver=2&hsh=4&fclid=2c0b3cdd-6692-699b-18bd-2bc9674768a3&u=a1aHR0cHM6Ly93d3cuemhpaHUuY29tL3F1ZXN0aW9uLzExNTA1Njc3MDI4&ntb=1
SDXL、FLUX和Pony三个模型在技术架构、应用场景和性能特点上各有不同,以下是它们的对比分析: 技术架构 SDXL:基于Stable Diffusion架构,属于通用图像生成模型,支持多种风格和高质量图像生 …
arxiv.org/abs/1902.07432v1
Online social networks have become the medium for efficient viral marketing exploiting social influence in information diffusion. However, the emerging application Social Coupon (SC) incorporating social referral into coupons cannot be efficiently so...
arxiv.org/abs/2104.09311v2
We study finite-time horizon continuous-time linear-convex reinforcement learning problems in an episodic setting. In this problem, the unknown linear jump-diffusion process is controlled subject to nonsmooth convex costs. We show that the associated...
github.com/facebookresearch/DiT
Official PyTorch Implementation of "Scalable Diffusion Models with Transformers" (⭐ 8396)
arxiv.org/abs/2603.04893v1
Diverse outputs in text generation are necessary for effective exploration in complex reasoning tasks, such as code generation and mathematical problem solving. Such Pass@$k$ problems benefit from distinct candidates covering the solution space. Howe...
arxiv.org/abs/2411.04919v2
Visual imitation learning methods demonstrate strong performance, yet they lack generalization when faced with visual input perturbations, including variations in lighting and textures, impeding their real-world application. We propose Stem-OB that u...
arxiv.org/abs/2012.12319v1
Bile, the central metabolic product of the liver, is secreted by hepatocytes into bile canaliculi (BC), tubular subcellular structures of 0.5-2 $μ$m diameter which are formed by the apical membranes of juxtaposed hepatocytes. BC interconnect to buil...
github.com/tencent-ailab/SongBloom
The official code repository for SongBloom: Coherent Song Generation via Interleaved Autoregressive Sketching and Diffusion Refinement (⭐ 759)
arxiv.org/abs/2512.18254v1
Interleaved text-image generation aims to jointly produce coherent visual frames and aligned textual descriptions within a single sequence, enabling tasks such as style transfer, compositional synthesis, and procedural tutorials. We present Loom, a u...
arxiv.org/abs/0807.0721v1
The paper shows how the diffusive movement of ions through a channel protein can be described as a chemical reaction over an arbitrary shaped potential barrier. The result is simple and intuitive but without approximation beyond the electrodiffusio...
arxiv.org/abs/2112.08739v3
The widespread diffusion of synthetically generated content is a serious threat that needs urgent countermeasures. As a matter of fact, the generation of synthetic content is not restricted to multimedia data like videos, photographs or audio sequenc...
github.com/TencentQQGYLab/ELLA
ELLA: Equip Diffusion Models with LLM for Enhanced Semantic Alignment (⭐ 1279)
arxiv.org/abs/2407.13752v1
Recent advances in text-to-image model customization have underscored the importance of integrating new concepts with a few examples. Yet, these progresses are largely confined to widely recognized subjects, which can be learned with relative ease th...
github.com/prophesier/diff-svc
Singing Voice Conversion via diffusion model (⭐ 2720)
github.com/n00mkrad/text2image-gui
Somewhat modular text2image GUI, initially just for Stable Diffusion (⭐ 967)
arxiv.org/abs/2406.04769v2
Field-of-view (FOV) recovery of truncated chest CT scans is crucial for accurate body composition analysis, which involves quantifying skeletal muscle and subcutaneous adipose tissue (SAT) on CT slices. This, in turn, enables disease prognostication....
arxiv.org/abs/2102.06942v1
Convolutional networks are successful, but they have recently been outperformed by new neural networks that are equivariant under rotations and translations. These new networks work better because they do not struggle with learning each possible orie...
arxiv.org/abs/0812.3083v1
In the present paper we present a finite element approach for option pricing in the framework of a well-known stochastic volatility model with jumps, the Bates model. In this model the asset log-returns are assumed to follow a jump-diffusion model...