mirror of
https://github.com/NVIDIA/TensorRT-LLM.git
synced 2026-01-14 06:27:45 +08:00
* Feat: Offload ptable to cpu if enable_chunk_context Signed-off-by: Kate Cheng <yunhsuanc@nvidia.com> * Feat: offload ptable to cpu for chunk context mode Signed-off-by: Kate Cheng <yunhsuanc@nvidia.com> * Fix and add comment Signed-off-by: Kate Cheng <yunhsuanc@nvidia.com> * Update Readme for multimodal and add a new param mm_embedding_offloading Signed-off-by: Kate Cheng <yunhsuanc@nvidia.com> * fix: Correct prompt table offloading condition in PromptTuningBuffers Signed-off-by: Kate Cheng <yunhsuanc@nvidia.com> * Clean up the code Signed-off-by: Kate Cheng <yunhsuanc@nvidia.com> * Add commits to explain copy from cpu <-> gpu using pinned memory Signed-off-by: Kate Cheng <yunhsuanc@nvidia.com> * Fix namings based on comments Signed-off-by: Kate Cheng <yunhsuanc@nvidia.com> * Fix format based on precommit Signed-off-by: Kate Cheng <yunhsuanc@nvidia.com> * Modify --mm_embedding_offloading flag Signed-off-by: Kate Cheng <yunhsuanc@nvidia.com> --------- Signed-off-by: Kate Cheng <yunhsuanc@nvidia.com> Co-authored-by: Haohang Huang <31998628+symphonylyh@users.noreply.github.com> |
||
|---|---|---|
| .. | ||
| allocateKvCache.h | ||
| assignReqSeqSlots.h | ||
| cacheTransceiver.h | ||
| capacityScheduler.h | ||
| common.h | ||
| contextProgress.h | ||
| createNewDecoderRequests.h | ||
| decoderBuffers.h | ||
| evictionPolicy.h | ||
| generateRequestOptions.h | ||
| guidedDecoder.h | ||
| handleContextLogits.h | ||
| handleGenerationLogits.h | ||
| kvCacheConfig.h | ||
| kvCacheEventManager.h | ||
| kvCacheManager.h | ||
| kvCacheTransferManager.h | ||
| kvCacheUtils.h | ||
| llmRequest.h | ||
| logitsPostProcessor.h | ||
| makeDecodingBatchInputOutput.h | ||
| medusaBuffers.h | ||
| microBatchScheduler.h | ||
| pauseRequests.h | ||
| peftCacheManager.h | ||
| peftCacheManagerConfig.h | ||
| promptTuningBuffers.h | ||
| rnnStateManager.h | ||
| runtimeBuffers.h | ||
| sequenceSlotManager.h | ||
| transformerBuffers.h | ||
| trtGptModelOptionalParams.h | ||