| .. |
|
cache_transmission
|
[None][feat] Support for cancelling requests with disaggregation (#8114)
|
2025-10-02 11:04:26 -07:00 |
|
cacheTransceiverConfig.cpp
|
chore:[BREAKING CHANGE] use cacheTransceiverConfig as knobs for disagg service (#5234)
|
2025-07-17 17:42:07 +08:00 |
|
CMakeLists.txt
|
Cache transceiver support VSWA (#5505)
|
2025-07-05 01:18:42 +09:00 |
|
contextPhaseParams.cpp
|
Update TensorRT-LLM (#2936)
|
2025-03-18 21:25:19 +08:00 |
|
debugConfig.cpp
|
Update TensorRT-LLM (#2873)
|
2025-03-11 21:13:42 +08:00 |
|
decodingConfig.cpp
|
Feat: Variable-Beam-Width-Search (VBWS) part3 (#3338)
|
2025-04-08 23:51:27 +08:00 |
|
disaggServerUtil.cpp
|
Update TensorRT-LLM (#2873)
|
2025-03-11 21:13:42 +08:00 |
|
dynamicBatchConfig.cpp
|
Update TensorRT-LLM (#2873)
|
2025-03-11 21:13:42 +08:00 |
|
dynamicBatchTuner.cpp
|
Update TensorRT-LLM (#2873)
|
2025-03-11 21:13:42 +08:00 |
|
dynamicBatchTuner.h
|
Update TensorRT-LLM (#2873)
|
2025-03-11 21:13:42 +08:00 |
|
executor.cpp
|
[TRTLLM-6881][feat] Include attention dp rank info with KV cache events (#6563)
|
2025-08-07 14:17:07 +02:00 |
|
executorConfig.cpp
|
[nvbug/5374773] chore: Add a runtime flag to enable fail fast when attn window is too large to fit at least one sequence in KV cache (#5974)
|
2025-07-25 18:10:40 -04:00 |
|
executorImpl.cpp
|
[TRTLLM-6106][feat] Add support for KVCache transfer from KVCache reuse path (#6348)
|
2025-09-27 19:29:30 -04:00 |
|
executorImpl.h
|
[TRTLLM-6106][feat] Add support for KVCache transfer from KVCache reuse path (#6348)
|
2025-09-27 19:29:30 -04:00 |
|
executorKVCacheEventManager.cpp
|
Update TensorRT-LLM (#2873)
|
2025-03-11 21:13:42 +08:00 |
|
extendedRuntimePerfKnobConfig.cpp
|
Update TensorRT-LLM (#2873)
|
2025-03-11 21:13:42 +08:00 |
|
guidedDecodingConfig.cpp
|
Update TensorRT-LLM (#2873)
|
2025-03-11 21:13:42 +08:00 |
|
guidedDecodingParams.cpp
|
Update TensorRT-LLM (#2873)
|
2025-03-11 21:13:42 +08:00 |
|
intervalSet.h
|
Update TensorRT-LLM (#2873)
|
2025-03-11 21:13:42 +08:00 |
|
jsonSerialization.cpp
|
[TRTLLM-5000][feat] NGrams V2 (#4569)
|
2025-06-27 23:00:17 +08:00 |
|
kvCacheConfig.cpp
|
fix/improve kvcache allocation in PyTorch runtime (#5933)
|
2025-08-26 12:40:22 +08:00 |
|
kvCacheRetentionConfig.cpp
|
[None][feat] Nixl support for GDS (#5488)
|
2025-09-09 13:00:38 +08:00 |
|
logitsPostProcessorConfig.cpp
|
Update TensorRT-LLM (#2873)
|
2025-03-11 21:13:42 +08:00 |
|
loraConfig.cpp
|
[TRTLLM-6683][feat] Support LoRA reload CPU cache evicted adapter (#6510)
|
2025-08-07 09:05:36 +03:00 |
|
model.h
|
fix: request termination in pipeline parallelism (#3892)
|
2025-05-05 21:51:41 +08:00 |
|
mropeConfig.cpp
|
Update TensorRT-LLM (#2873)
|
2025-03-11 21:13:42 +08:00 |
|
multimodalInput.cpp
|
[TRTLLM-5007][feat] Add multimodal hashing support (image hashing) (#4145)
|
2025-06-10 01:59:56 +08:00 |
|
orchestratorConfig.cpp
|
Update TensorRT-LLM (#2873)
|
2025-03-11 21:13:42 +08:00 |
|
orchestratorUtils.h
|
refactor: Introduce MpiTag enumeration and update MPI function signatures (#3893)
|
2025-05-04 13:24:29 +02:00 |
|
outputConfig.cpp
|
[TRTLLM-3429] feat: Overlap scheduling in C++ runtime (#3625)
|
2025-05-06 15:06:46 +02:00 |
|
parallelConfig.cpp
|
feat: Add numNodes to ParallelConfig (#3346)
|
2025-04-13 13:55:04 +02:00 |
|
peftCacheConfig.cpp
|
Update TensorRT-LLM (#2873)
|
2025-03-11 21:13:42 +08:00 |
|
promptTuningConfig.cpp
|
Update TensorRT-LLM (#2873)
|
2025-03-11 21:13:42 +08:00 |
|
request.cpp
|
[TRTLLM-7398][feat] Support KV cache salting for secure KV cache reuse (#7106)
|
2025-09-06 17:58:32 -04:00 |
|
requestImpl.h
|
[TRTLLM-7398][feat] Support KV cache salting for secure KV cache reuse (#7106)
|
2025-09-06 17:58:32 -04:00 |
|
requestUtils.cpp
|
Update TensorRT-LLM (#2873)
|
2025-03-11 21:13:42 +08:00 |
|
requestUtils.h
|
Update TensorRT-LLM (#2873)
|
2025-03-11 21:13:42 +08:00 |
|
requestWithId.cpp
|
Update TensorRT-LLM (#2873)
|
2025-03-11 21:13:42 +08:00 |
|
requestWithId.h
|
Update TensorRT-LLM (#2873)
|
2025-03-11 21:13:42 +08:00 |
|
response.cpp
|
Update TensorRT-LLM (#2873)
|
2025-03-11 21:13:42 +08:00 |
|
responseImpl.h
|
Update TensorRT-LLM (#2873)
|
2025-03-11 21:13:42 +08:00 |
|
samplingConfig.cpp
|
Feat: Variable-Beam-Width-Search (VBWS) part4 (#3979)
|
2025-05-12 22:32:29 +02:00 |
|
schedulerConfig.cpp
|
Update TensorRT-LLM (#2873)
|
2025-03-11 21:13:42 +08:00 |
|
serialization.cpp
|
[TRTLLM-6106][feat] Add support for KVCache transfer from KVCache reuse path (#6348)
|
2025-09-27 19:29:30 -04:00 |
|
serializeUtils.h
|
[TRTLLM-6106][feat] Add support for KVCache transfer from KVCache reuse path (#6348)
|
2025-09-27 19:29:30 -04:00 |
|
speculativeDecodingConfig.cpp
|
Update TensorRT-LLM (#2873)
|
2025-03-11 21:13:42 +08:00 |
|
tensor.cpp
|
[TRTLLM-4629] [feat] Add support of CUDA13 and sm103 devices (#7568)
|
2025-09-16 09:56:18 +08:00 |
|
types.cpp
|
Update TensorRT-LLM (#2873)
|
2025-03-11 21:13:42 +08:00 |