..
cache_transmission
[None][feat] Add Request specific exception ( #6931 )
2025-09-04 18:43:42 -04:00
cacheTransceiverConfig.cpp
chore:[BREAKING CHANGE] use cacheTransceiverConfig as knobs for disagg service ( #5234 )
2025-07-17 17:42:07 +08:00
CMakeLists.txt
Cache transceiver support VSWA ( #5505 )
2025-07-05 01:18:42 +09:00
contextPhaseParams.cpp
Update TensorRT-LLM ( #2936 )
2025-03-18 21:25:19 +08:00
debugConfig.cpp
Update TensorRT-LLM ( #2873 )
2025-03-11 21:13:42 +08:00
decodingConfig.cpp
Feat: Variable-Beam-Width-Search (VBWS) part3 ( #3338 )
2025-04-08 23:51:27 +08:00
disaggServerUtil.cpp
Update TensorRT-LLM ( #2873 )
2025-03-11 21:13:42 +08:00
dynamicBatchConfig.cpp
Update TensorRT-LLM ( #2873 )
2025-03-11 21:13:42 +08:00
dynamicBatchTuner.cpp
Update TensorRT-LLM ( #2873 )
2025-03-11 21:13:42 +08:00
dynamicBatchTuner.h
Update TensorRT-LLM ( #2873 )
2025-03-11 21:13:42 +08:00
executor.cpp
[TRTLLM-6881][feat] Include attention dp rank info with KV cache events ( #6563 )
2025-08-07 14:17:07 +02:00
executorConfig.cpp
[nvbug/5374773] chore: Add a runtime flag to enable fail fast when attn window is too large to fit at least one sequence in KV cache ( #5974 )
2025-07-25 18:10:40 -04:00
executorImpl.cpp
feat: Support structural tag in C++ runtime and upgrade xgrammar to 0.1.21 ( #6408 )
2025-07-31 09:53:52 +08:00
executorImpl.h
chore: Mass integration of release/0.20 ( #5082 )
2025-06-17 14:32:02 +03:00
executorKVCacheEventManager.cpp
Update TensorRT-LLM ( #2873 )
2025-03-11 21:13:42 +08:00
extendedRuntimePerfKnobConfig.cpp
Update TensorRT-LLM ( #2873 )
2025-03-11 21:13:42 +08:00
guidedDecodingConfig.cpp
Update TensorRT-LLM ( #2873 )
2025-03-11 21:13:42 +08:00
guidedDecodingParams.cpp
Update TensorRT-LLM ( #2873 )
2025-03-11 21:13:42 +08:00
intervalSet.h
Update TensorRT-LLM ( #2873 )
2025-03-11 21:13:42 +08:00
jsonSerialization.cpp
[TRTLLM-5000][feat] NGrams V2 ( #4569 )
2025-06-27 23:00:17 +08:00
kvCacheConfig.cpp
fix/improve kvcache allocation in PyTorch runtime ( #5933 )
2025-08-26 12:40:22 +08:00
kvCacheRetentionConfig.cpp
feature: KV Cache GPUDirect Storage ( #3209 )
2025-05-28 23:27:43 +00:00
logitsPostProcessorConfig.cpp
Update TensorRT-LLM ( #2873 )
2025-03-11 21:13:42 +08:00
loraConfig.cpp
[TRTLLM-6683][feat] Support LoRA reload CPU cache evicted adapter ( #6510 )
2025-08-07 09:05:36 +03:00
model.h
fix: request termination in pipeline parallelism ( #3892 )
2025-05-05 21:51:41 +08:00
mropeConfig.cpp
Update TensorRT-LLM ( #2873 )
2025-03-11 21:13:42 +08:00
multimodalInput.cpp
[TRTLLM-5007][feat] Add multimodal hashing support (image hashing) ( #4145 )
2025-06-10 01:59:56 +08:00
orchestratorConfig.cpp
Update TensorRT-LLM ( #2873 )
2025-03-11 21:13:42 +08:00
orchestratorUtils.h
refactor: Introduce MpiTag enumeration and update MPI function signatures ( #3893 )
2025-05-04 13:24:29 +02:00
outputConfig.cpp
[TRTLLM-3429] feat: Overlap scheduling in C++ runtime ( #3625 )
2025-05-06 15:06:46 +02:00
parallelConfig.cpp
feat: Add numNodes to ParallelConfig ( #3346 )
2025-04-13 13:55:04 +02:00
peftCacheConfig.cpp
Update TensorRT-LLM ( #2873 )
2025-03-11 21:13:42 +08:00
promptTuningConfig.cpp
Update TensorRT-LLM ( #2873 )
2025-03-11 21:13:42 +08:00
request.cpp
[TRTLLM-7398][feat] Support KV cache salting for secure KV cache reuse ( #7106 )
2025-09-06 17:58:32 -04:00
requestImpl.h
[TRTLLM-7398][feat] Support KV cache salting for secure KV cache reuse ( #7106 )
2025-09-06 17:58:32 -04:00
requestUtils.cpp
Update TensorRT-LLM ( #2873 )
2025-03-11 21:13:42 +08:00
requestUtils.h
Update TensorRT-LLM ( #2873 )
2025-03-11 21:13:42 +08:00
requestWithId.cpp
Update TensorRT-LLM ( #2873 )
2025-03-11 21:13:42 +08:00
requestWithId.h
Update TensorRT-LLM ( #2873 )
2025-03-11 21:13:42 +08:00
response.cpp
Update TensorRT-LLM ( #2873 )
2025-03-11 21:13:42 +08:00
responseImpl.h
Update TensorRT-LLM ( #2873 )
2025-03-11 21:13:42 +08:00
samplingConfig.cpp
Feat: Variable-Beam-Width-Search (VBWS) part4 ( #3979 )
2025-05-12 22:32:29 +02:00
schedulerConfig.cpp
Update TensorRT-LLM ( #2873 )
2025-03-11 21:13:42 +08:00
serialization.cpp
[TRTLLM-7398][feat] Support KV cache salting for secure KV cache reuse ( #7106 )
2025-09-06 17:58:32 -04:00
serializeUtils.h
[TRTLLM-6881][feat] Include attention dp rank info with KV cache events ( #6563 )
2025-08-07 14:17:07 +02:00
speculativeDecodingConfig.cpp
Update TensorRT-LLM ( #2873 )
2025-03-11 21:13:42 +08:00
tensor.cpp
Update TensorRT-LLM ( #2873 )
2025-03-11 21:13:42 +08:00
types.cpp
Update TensorRT-LLM ( #2873 )
2025-03-11 21:13:42 +08:00