TensorRT-LLMs

mirror of https://github.com/NVIDIA/TensorRT-LLM.git synced 2026-01-14 06:27:45 +08:00

History

Chuang Zhu ffc0b8f5da Cache transceiver support VSWA (#5505 ) Signed-off-by: ShiXiaowei02 <39303645+Shixiaowei02@users.noreply.github.com> Signed-off-by: Chuang Zhu <111838961+chuangz0@users.noreply.github.com> Co-authored-by: ShiXiaowei02 <39303645+Shixiaowei02@users.noreply.github.com>		2025-07-05 01:18:42 +09:00
..
cache_transmission	Cache transceiver support VSWA (#5505 )	2025-07-05 01:18:42 +09:00
cacheTransceiverConfig.cpp	[TRTLLM-3429] feat: Overlap scheduling in C++ runtime (#3625 )	2025-05-06 15:06:46 +02:00
CMakeLists.txt	Cache transceiver support VSWA (#5505 )	2025-07-05 01:18:42 +09:00
contextPhaseParams.cpp	Update TensorRT-LLM (#2936 )	2025-03-18 21:25:19 +08:00
debugConfig.cpp	Update TensorRT-LLM (#2873 )	2025-03-11 21:13:42 +08:00
decodingConfig.cpp	Feat: Variable-Beam-Width-Search (VBWS) part3 (#3338 )	2025-04-08 23:51:27 +08:00
disaggServerUtil.cpp	Update TensorRT-LLM (#2873 )	2025-03-11 21:13:42 +08:00
dynamicBatchConfig.cpp	Update TensorRT-LLM (#2873 )	2025-03-11 21:13:42 +08:00
dynamicBatchTuner.cpp	Update TensorRT-LLM (#2873 )	2025-03-11 21:13:42 +08:00
dynamicBatchTuner.h	Update TensorRT-LLM (#2873 )	2025-03-11 21:13:42 +08:00
executor.cpp	Update TensorRT-LLM (#2873 )	2025-03-11 21:13:42 +08:00
executorConfig.cpp	Feat: Variable-Beam-Width-Search (VBWS) part4 (#3979 )	2025-05-12 22:32:29 +02:00
executorImpl.cpp	test: Add LLGuidance test and refine guided decoding (#5348 )	2025-06-25 14:12:56 +08:00
executorImpl.h	chore: Mass integration of release/0.20 (#5082 )	2025-06-17 14:32:02 +03:00
executorKVCacheEventManager.cpp	Update TensorRT-LLM (#2873 )	2025-03-11 21:13:42 +08:00
extendedRuntimePerfKnobConfig.cpp	Update TensorRT-LLM (#2873 )	2025-03-11 21:13:42 +08:00
guidedDecodingConfig.cpp	Update TensorRT-LLM (#2873 )	2025-03-11 21:13:42 +08:00
guidedDecodingParams.cpp	Update TensorRT-LLM (#2873 )	2025-03-11 21:13:42 +08:00
intervalSet.h	Update TensorRT-LLM (#2873 )	2025-03-11 21:13:42 +08:00
jsonSerialization.cpp	[TRTLLM-5000][feat] NGrams V2 (#4569 )	2025-06-27 23:00:17 +08:00
kvCacheConfig.cpp	refactor: remove batch_manager::KvCacheConfig and use executor::KvCacheConfig instead (#5384 )	2025-06-26 19:45:52 +08:00
kvCacheRetentionConfig.cpp	feature: KV Cache GPUDirect Storage (#3209 )	2025-05-28 23:27:43 +00:00
logitsPostProcessorConfig.cpp	Update TensorRT-LLM (#2873 )	2025-03-11 21:13:42 +08:00
loraConfig.cpp	Update TensorRT-LLM (#2873 )	2025-03-11 21:13:42 +08:00
model.h	fix: request termination in pipeline parallelism (#3892 )	2025-05-05 21:51:41 +08:00
mropeConfig.cpp	Update TensorRT-LLM (#2873 )	2025-03-11 21:13:42 +08:00
multimodalInput.cpp	[TRTLLM-5007][feat] Add multimodal hashing support (image hashing) (#4145 )	2025-06-10 01:59:56 +08:00
orchestratorConfig.cpp	Update TensorRT-LLM (#2873 )	2025-03-11 21:13:42 +08:00
orchestratorUtils.h	refactor: Introduce MpiTag enumeration and update MPI function signatures (#3893 )	2025-05-04 13:24:29 +02:00
outputConfig.cpp	[TRTLLM-3429] feat: Overlap scheduling in C++ runtime (#3625 )	2025-05-06 15:06:46 +02:00
parallelConfig.cpp	feat: Add numNodes to ParallelConfig (#3346 )	2025-04-13 13:55:04 +02:00
peftCacheConfig.cpp	Update TensorRT-LLM (#2873 )	2025-03-11 21:13:42 +08:00
promptTuningConfig.cpp	Update TensorRT-LLM (#2873 )	2025-03-11 21:13:42 +08:00
request.cpp	[TRTLLM-5007][feat] Add multimodal hashing support (image hashing) (#4145 )	2025-06-10 01:59:56 +08:00
requestImpl.h	[TRTLLM-5007][feat] Add multimodal hashing support (image hashing) (#4145 )	2025-06-10 01:59:56 +08:00
requestUtils.cpp	Update TensorRT-LLM (#2873 )	2025-03-11 21:13:42 +08:00
requestUtils.h	Update TensorRT-LLM (#2873 )	2025-03-11 21:13:42 +08:00
requestWithId.cpp	Update TensorRT-LLM (#2873 )	2025-03-11 21:13:42 +08:00
requestWithId.h	Update TensorRT-LLM (#2873 )	2025-03-11 21:13:42 +08:00
response.cpp	Update TensorRT-LLM (#2873 )	2025-03-11 21:13:42 +08:00
responseImpl.h	Update TensorRT-LLM (#2873 )	2025-03-11 21:13:42 +08:00
samplingConfig.cpp	Feat: Variable-Beam-Width-Search (VBWS) part4 (#3979 )	2025-05-12 22:32:29 +02:00
schedulerConfig.cpp	Update TensorRT-LLM (#2873 )	2025-03-11 21:13:42 +08:00
serialization.cpp	[TRTLLM-5000][feat] NGrams V2 (#4569 )	2025-06-27 23:00:17 +08:00
serializeUtils.h	[TRTLLM-5007][feat] Add multimodal hashing support (image hashing) (#4145 )	2025-06-10 01:59:56 +08:00
speculativeDecodingConfig.cpp	Update TensorRT-LLM (#2873 )	2025-03-11 21:13:42 +08:00
tensor.cpp	Update TensorRT-LLM (#2873 )	2025-03-11 21:13:42 +08:00
types.cpp	Update TensorRT-LLM (#2873 )	2025-03-11 21:13:42 +08:00