TensorRT-LLMs

mirror of https://github.com/NVIDIA/TensorRT-LLM.git synced 2026-01-24 04:33:04 +08:00

History

liji-nv dca6397d1e feat: Introduce UB allocator for pytorch flow (#3257 ) * Instead of allocating UserBuffers at beginning of runtime, UB buffers are now managed with global allocator. The allocator will dynamically assign free UB buffer or allocate new buffer for torch tensor. It makes userbuffers easier to use. * In common usecase, the Userbuffers will be allocated correctly during warm up stage. There is no dynamic allocation during inference. * UB fusion pattern is rewroten using the new UB Allocator. It contains following passes: 1. Fuse Quant with allreduce, replace with UB impl, and insert a copy_to_userbuffers. Currently the normal allreduce still does not support FP8 quant. So this need to be done in UB pass 2. Convert all supported allreduce with UB and insert copy_to_userbuffers. 3. Fuse op before ar with the copy_to_userbuffers. So the op directly writes to the userbuffer 4. Remove userbuffers finalize if the output is connect to another UB allreduce. Signed-off-by: Jin Li <59594262+liji-nv@users.noreply.github.com>		2025-04-08 18:39:49 +08:00
..
batch_manager	feat: Support PeftCacheManager in Torch (#3186 )	2025-04-04 12:38:08 +08:00
common	feat: support abort disconnected requests (#3214 )	2025-04-07 16:14:58 +08:00
executor	feat: Add option to run disaggregated serving without ctx servers,… (#3243 )	2025-04-07 21:56:03 -04:00
runtime	refactor: Expose DecoderState via bindings and integrate in TRTLLMDecoder (#3139 )	2025-04-05 07:42:35 +08:00
userbuffers	feat: Introduce UB allocator for pytorch flow (#3257 )	2025-04-08 18:39:49 +08:00
bindings.cpp	feat: Support PeftCacheManager in Torch (#3186 )	2025-04-04 12:38:08 +08:00
CMakeLists.txt	Update TensorRT-LLM (#2849 )	2025-03-04 18:44:00 +08:00