130 results for GPUs

news.ycombinator.com
Show HN: Loft CLI – Fine-tune and run LLMs (1–3B) on 8 GB MacBook Air, no GPUshttps://news.ycombinator.com/item?id=44627754GitHub: https://github.com/diptanshu1991/LoFTI built *LoFT*, a lightweight CLI that turns any 8 GB laptop into a tiny LLM training and inference rig — no GPU, no cloud.5 Commands: 1. `loft finetune` → Train LoRA adapters on CPU 2. `loft merge` → Merge adapters into model 3. `loft export` → Convert to GGUF (FP16) 4. `loft quantize` → Apply Q4_0 (4-bit) quantization 5. `loft chat` → llama.cpp CPU chat @ ~7 tok/sBenchmarks on 8 GB MacBook Air: | Step | Time | Peak RAM | |-------------|--------|----------| | Finetune | 23 min (sample run) | 308 MB | | Merge | 4.7 min | 322 MB | | Quantize | 21 sec | 322 MB | | Inference | 6.9 tok/s | 322 MB |Also ran a full 300-row Dolly finetune (2 epochs) in *~1.5 hours*, achieving *sub-1 loss* on CPU-only setup. No crashes, swap kills, or GPU needed.Why this matters: - Makes local LLM customization accessible to devs without GPU access - Enables domain-specific agents (summarizer, support bot, Q&A) on commodity laptops - Everything runs via CPU (no CUDA
Jul 20, 2025 6:14 PM
bing.com
Google introduces Colab Pro w/ faster GPUs, more memory, and longer runtimeshttp://www.bing.com/news/apiclick.aspx?ref=FexRss&aid=&tid=6a8d561189d34999b4d957ebb47ae060&url=https%3A%2F%2F9to5google.com%2F2020%2F02%2F08%2Fgoogle-introduces-colab-pro%2F&c=12203660269233921366&mkt=en-usGoogle Colab is a useful tool for data scientists and AI researchers sharing work online. The company this week quietly introduced a paid “Colab Pro” tier with three benefits. Within a Colaboratory’s ...
Feb 8, 2020 9:30 AM
twitter.com
Show HN: GPU Accelerated PDALhttps://twitter.com/ZyMazza/status/2092260677301817482I made a fork of PDAL with GPU acceleration for NVIDIA GPUs! On common pipelines, you can get anywhere from 2 to 14x speed improvements over stock PDAL. It's a drop in replacement, just use npm i gpupdal to get it and use gpupdal for any command where you would use pdal.(PDAL is an open source library for working with point cloud data, for the uninitiated)Full source code here: https://github.com/zymazza/GPUPDAL
Aug 25, 2026 3:20 PM
bing.com
RAM buyers are competing with bots nearly 10 to 1http://www.bing.com/news/apiclick.aspx?ref=FexRss&aid=&tid=6a8cfdcd808348bc8cc8547feb26847d&url=https%3A%2F%2Fwww.msn.com%2Fen-us%2Fnews%2Fother%2Fram-buyers-are-competing-with-bots-nearly-10-to-1%2Far-AA2aPejK&c=7760214341597906176&mkt=en-usSo it went with GPUs and webcams and toilet paper in the pandemic, so it goes with RAM in 2026. For at least one retailer, over 90% of the web traffic for RAM listings came from “scalper” bots. That’s ...
Aug 24, 2026 9:19 AM
github.com
Show HN: Icebug-format: immutable, interoperable graph standardhttps://github.com/Ladybug-Memory/icebug-formatMost graph analytics packages have a mutable graph implementation that uses a heap allocated vector to store the graph. It works for toy graphs. But if you're loading a billion edge graph using G.add_edge() it's going to take a while.We don't need to invent new standards. Such interoperable, immutable memory standards already exist: Apache Arrow and Compressed Sparse Rows (CSR). CSR is widely used in scipy, cugraph and columnar graph databases among others. Both on CPUs and GPUs.icebug-format combines both into a on-disk standard based on Apache Parquet and an in-memory format based on Apache Arrow.Bindings available in many popular languages including python, typescript and rust.The package ships with convenience scripts to convert flat tables such as vertex.parquet and edges.parquet to this format in RAM/disk constrained environments.Sample graphs: https://huggingface.co/datasets/ladybugdb/ldbc-csr/tree/mainConverted from: https://ldbcouncil.org/benchmarks/graphalytics/datasets/Large
Aug 20, 2026 11:57 PM
techcrunch.com
Meet the startup helping Wall Street put a price on AI computehttps://techcrunch.com/video/meet-the-startup-helping-wall-street-put-a-price-on-ai-compute/The AI buildout shows no signs of slowing. And with hundreds of billions of dollars a year going into data centers and GPUs, compute has become the single biggest cost for anyone building AI products. But for all that spending, there still isn’t a straightforward way to put a price on compute — or for firms to hedge their exposure when the price changes.  Silicon Data […]
Aug 19, 2026 5:26 PM
bing.com
The Future Of AI Compute Won’t Run On Just One Kind Of Chiphttp://www.bing.com/news/apiclick.aspx&aid=&tid=6a85ce0fc09f41dc82d0f81200c65eea&url=https%3a%2f%2fsemiengineering.com%2fthe-future-of-ai-compute-wont-run-on-just-one-kind-of-chip%2f&c=14211995415715052224&mkt=en-usPower, token cost, interconnects, and software orchestration are pushing data centers toward heterogeneous clusters built from CPUs, GPUs, NPUs, optics, and custom accelerators.
Aug 19, 2026 12:01 AM
bing.com
The Future Of AI Compute Won’t Run On Just One Kind Of Chiphttp://www.bing.com/news/apiclick.aspx&aid=&tid=6a85ebb55f7b463ca3a58e7b8c30bea3&url=https%3a%2f%2fsemiengineering.com%2fthe-future-of-ai-compute-wont-run-on-just-one-kind-of-chip%2f&c=14211995415715052224&mkt=en-usPower, token cost, interconnects, and software orchestration are pushing data centers toward heterogeneous clusters built from CPUs, GPUs, NPUs, optics, and custom accelerators.
Aug 19, 2026 12:01 AM
bing.com
The Future Of AI Compute Won’t Run On Just One Kind Of Chiphttp://www.bing.com/news/apiclick.aspx&aid=&tid=6a85e10917f5482781e74d821cc7375e&url=https%3a%2f%2fsemiengineering.com%2fthe-future-of-ai-compute-wont-run-on-just-one-kind-of-chip%2f&c=14211995415715052224&mkt=en-usPower, token cost, interconnects, and software orchestration are pushing data centers toward heterogeneous clusters built from CPUs, GPUs, NPUs, optics, and custom accelerators.
Aug 19, 2026 12:01 AM
bing.com
The Future Of AI Compute Won’t Run On Just One Kind Of Chiphttp://www.bing.com/news/apiclick.aspx&aid=&tid=6a85d0359c1c41ddaca0847f3f803084&url=https%3a%2f%2fsemiengineering.com%2fthe-future-of-ai-compute-wont-run-on-just-one-kind-of-chip%2f&c=14211995415715052224&mkt=en-usPower, token cost, interconnects, and software orchestration are pushing data centers toward heterogeneous clusters built from CPUs, GPUs, NPUs, optics, and custom accelerators.
Aug 19, 2026 12:01 AM
bing.com
The Future Of AI Compute Won’t Run On Just One Kind Of Chiphttp://www.bing.com/news/apiclick.aspx&aid=&tid=6a85dab649a84e5cae378f9ecc00cd2b&url=https%3a%2f%2fsemiengineering.com%2fthe-future-of-ai-compute-wont-run-on-just-one-kind-of-chip%2f&c=14211995415715052224&mkt=en-usPower, token cost, interconnects, and software orchestration are pushing data centers toward heterogeneous clusters built from CPUs, GPUs, NPUs, optics, and custom accelerators.
Aug 19, 2026 12:01 AM
bing.com
The Future Of AI Compute Won’t Run On Just One Kind Of Chiphttp://www.bing.com/news/apiclick.aspx&aid=&tid=6a85e547f87f485299a3d708ab8028d4&url=https%3a%2f%2fsemiengineering.com%2fthe-future-of-ai-compute-wont-run-on-just-one-kind-of-chip%2f&c=14211995415715052224&mkt=en-usPower, token cost, interconnects, and software orchestration are pushing data centers toward heterogeneous clusters built from CPUs, GPUs, NPUs, optics, and custom accelerators.
Aug 19, 2026 12:01 AM
bing.com
The Future Of AI Compute Won’t Run On Just One Kind Of Chiphttp://www.bing.com/news/apiclick.aspx&aid=&tid=6a85d49a09974134b158554ff2e4ca8d&url=https%3a%2f%2fsemiengineering.com%2fthe-future-of-ai-compute-wont-run-on-just-one-kind-of-chip%2f&c=14211995415715052224&mkt=en-usPower, token cost, interconnects, and software orchestration are pushing data centers toward heterogeneous clusters built from CPUs, GPUs, NPUs, optics, and custom accelerators.
Aug 19, 2026 12:01 AM
bing.com
The Future Of AI Compute Won’t Run On Just One Kind Of Chiphttp://www.bing.com/news/apiclick.aspx&aid=&tid=6a85cc20cf624d94928113555e1d8b64&url=https%3a%2f%2fsemiengineering.com%2fthe-future-of-ai-compute-wont-run-on-just-one-kind-of-chip%2f&c=14211995415715052224&mkt=en-usPower, token cost, interconnects, and software orchestration are pushing data centers toward heterogeneous clusters built from CPUs, GPUs, NPUs, optics, and custom accelerators.
Aug 19, 2026 12:01 AM
bing.com
The Future Of AI Compute Won’t Run On Just One Kind Of Chiphttp://www.bing.com/news/apiclick.aspx&aid=&tid=6a85e773ad7f4f3cafb4f394bee1ad5c&url=https%3a%2f%2fsemiengineering.com%2fthe-future-of-ai-compute-wont-run-on-just-one-kind-of-chip%2f&c=14211995415715052224&mkt=en-usPower, token cost, interconnects, and software orchestration are pushing data centers toward heterogeneous clusters built from CPUs, GPUs, NPUs, optics, and custom accelerators.
Aug 19, 2026 12:01 AM
bing.com
The Future Of AI Compute Won’t Run On Just One Kind Of Chiphttp://www.bing.com/news/apiclick.aspx&aid=&tid=6a85d8a8b578449fbba6ba523481e0ea&url=https%3a%2f%2fsemiengineering.com%2fthe-future-of-ai-compute-wont-run-on-just-one-kind-of-chip%2f&c=14211995415715052224&mkt=en-usPower, token cost, interconnects, and software orchestration are pushing data centers toward heterogeneous clusters built from CPUs, GPUs, NPUs, optics, and custom accelerators.
Aug 19, 2026 12:01 AM
bing.com
The Future Of AI Compute Won’t Run On Just One Kind Of Chiphttp://www.bing.com/news/apiclick.aspx&aid=&tid=6a85d6ad52494a989318706de939ec9d&url=https%3a%2f%2fsemiengineering.com%2fthe-future-of-ai-compute-wont-run-on-just-one-kind-of-chip%2f&c=14211995415715052224&mkt=en-usPower, token cost, interconnects, and software orchestration are pushing data centers toward heterogeneous clusters built from CPUs, GPUs, NPUs, optics, and custom accelerators.
Aug 19, 2026 12:01 AM
bing.com
The Future Of AI Compute Won’t Run On Just One Kind Of Chiphttp://www.bing.com/news/apiclick.aspx&aid=&tid=6a85d26689d846d0a3e324dee4308989&url=https%3a%2f%2fsemiengineering.com%2fthe-future-of-ai-compute-wont-run-on-just-one-kind-of-chip%2f&c=14211995415715052224&mkt=en-usPower, token cost, interconnects, and software orchestration are pushing data centers toward heterogeneous clusters built from CPUs, GPUs, NPUs, optics, and custom accelerators.
Aug 19, 2026 12:01 AM
pantheongpu.com
Show HN: PantheonGPU – GPU health testing and AI workload benchmarkinghttps://pantheongpu.com/Hi HN, I built PantheonGPU because I wanted a better way to answer a simple question: is this GPU actually healthy and performing the way it should?A GPU can show normal temperatures and utilization and still be underperforming, unstable under certain workloads, or have memory, PCIe, or configuration issues.PantheonGPU actively tests the GPU instead of only monitoring telemetry. It currently includes 45+ tests covering compute, tensor workloads, memory, cache, PCIe, thermals, stability, and AI/LLM inference.It supports both NVIDIA CUDA and AMD ROCm.I’m also exploring a larger use case: running Pantheon across GPU fleets to identify individual GPUs that behave differently from the rest of a server or cluster.I’d especially appreciate feedback from people running AI infrastructure, multi-GPU systems, local LLMs, or GPU clouds.
Aug 18, 2026 6:47 PM