TensorRT-LLMs

mirror of https://github.com/NVIDIA/TensorRT-LLM.git synced 2026-01-14 06:27:45 +08:00

History

qixiang-99 bf4f7ad744 feat: add Pytorch support of Vision Encoder for multimodal models (#3791 ) * feat: Add rename_weights_with_regex function for dynamic weight key renaming Introduced a new utility function to rename weight keys in a dictionary based on regex pattern matching. This allows for flexible mapping of keys from Hugging Face naming conventions to TRT-LLM naming conventions, enhancing model compatibility and usability. Signed-off-by: qixiang-99 <203170375+qixiang-99@users.noreply.github.com> * feat: Implement SiglipVisionModel and related components Added the SiglipVisionModel along with its associated classes, including SiglipAttention, SiglipEncoderLayer, and SiglipEncoder. Additionally, a new test suite for the SiglipVisionModel has been created to ensure compatibility with Hugging Face outputs. Currently SiglipVisionModel support batch size larger than one. Also, inputs and outputs shape are same with the HF for compatibility. Signed-off-by: qixiang-99 <203170375+qixiang-99@users.noreply.github.com> * feat: Add CLIPVisionModel and associated components Introduced the CLIPVisionModel along with its related classes, including CLIPAttention, CLIPEncoderLayer, CLIPEncoder, and CLIPVisionTransformer. This implementation aligns with Hugging Face's CLIP architecture, ensuring compatibility in input and output shapes. Signed-off-by: qixiang-99 <203170375+qixiang-99@users.noreply.github.com> * feat: Enhance CLIPVisionModel with attention metadata preparation and unit tests Updated the CLIPVisionModel to include a method for preparing attention metadata, simplifying the model's usage. Additionally, added a comprehensive unit test suite for the CLIPVisionModel, ensuring compatibility with Hugging Face outputs and validating model performance across various scenarios. Signed-off-by: qixiang-99 <203170375+qixiang-99@users.noreply.github.com> * feat: Refactor SiglipVisionModel with attention metadata preparation and update unit tests Enhanced the SiglipVisionModel by adding a method to prepare attention metadata, streamlining its usage. Updated unit tests to validate model performance and compatibility with Hugging Face outputs, including adjustments to the configuration and test scenarios. Signed-off-by: qixiang-99 <203170375+qixiang-99@users.noreply.github.com> * refactor: Remove unused rotary_emb parameter from CLIP and Siglip attention classes Eliminated the rotary_emb parameter from the CLIPAttention and SiglipAttention classes to streamline the code. Updated unit tests to reflect changes in the model configurations, including clarifications in the default configurations sourced from Hugging Face. Signed-off-by: qixiang-99 <203170375+qixiang-99@users.noreply.github.com> * feat: Integrate CLIPVisionModel into LlavaNextInputProcessor and enhance weight loading Added CLIPVisionModel to the LlavaNextInputProcessor for improved vision processing. Updated the model loading mechanism to ensure compatibility with the new vision model and added attention metadata preparation. Removed debug print statements from weight renaming function for cleaner code. Signed-off-by: qixiang-99 <203170375+qixiang-99@users.noreply.github.com> * refactor: Remove unused max_position_embeddings from CLIPAttention and update Siglip classes to use CLIP components Removed the unused max_position_embeddings variable from the CLIPAttention class. Updated the Siglip classes to utilize CLIP components, specifically replacing SiglipEncoder and SiglipAttention with their CLIP counterparts, streamlining the codebase and enhancing consistency across models. Signed-off-by: qixiang-99 <203170375+qixiang-99@users.noreply.github.com> * refactor: Consolidate weight loading logic into a shared implementation Refactored the weight loading process across CLIP and Siglip models by using a new utility function, _load_weights_impl, to streamline the loading mechanism. This change enhances code maintainability and reduces redundancy in weight handling, ensuring consistent behavior across different model architectures. Signed-off-by: qixiang-99 <203170375+qixiang-99@users.noreply.github.com> * refactor: Simplify output handling in CLIP and Siglip models by removing output_hidden_states parameter Removed the output_hidden_states parameter from the CLIPEncoder and SiglipVisionTransformer classes, streamlining the output handling process. Updated the corresponding unit tests to reflect these changes and ensure compatibility with the new output structure. Signed-off-by: qixiang-99 <203170375+qixiang-99@users.noreply.github.com> * feat: Enhance LlavaNextInputProcessor with dynamic model loading and memory optimization Updated the LlavaNextInputProcessor to support dynamic model loading from local paths or Hugging Face, improving memory efficiency by partially loading the model components. Integrated the LlavaNextMultiModalProjector and adjusted weight loading to ensure compatibility with the new architecture. Signed-off-by: qixiang-99 <203170375+qixiang-99@users.noreply.github.com> --------- Signed-off-by: qixiang-99 <203170375+qixiang-99@users.noreply.github.com> Co-authored-by: Haohang Huang <31998628+symphonylyh@users.noreply.github.com>		2025-05-03 05:13:47 +08:00
..
_torch	feat: add Pytorch support of Vision Encoder for multimodal models (#3791 )	2025-05-03 05:13:47 +08:00
auto_parallel	chore: remove usernames from comments (#3291 )	2025-04-05 13:44:28 +08:00
bench	[feat]: Allow for a settable end-of-sequence/padding token in max throughput benchmark. (#3776 )	2025-05-01 09:42:46 +08:00
commands	Add smart router for moe (#3641 )	2025-04-23 12:21:59 +08:00
evaluate	[TRTLLM-4763][test] Accuracy test improvement (Part 3.6): Deprecate mmlu_llmapi.py (#3802 )	2025-04-23 23:05:13 +08:00
executor	feat: LogitsProcessor in PyTorch backend (#3145 )	2025-05-01 14:15:30 -07:00
inputs	feat: llama4 input processor (#3383 )	2025-04-25 16:47:14 -07:00
layers	feat: Add FP8 support for SM 120 (#3248 )	2025-04-14 16:05:41 -07:00
llmapi	feat: Support Top-K logprobs and prompt_logprobs in LLMAPI (#3388 )	2025-05-01 12:47:14 -04:00
models	chore: bump version to 0.19.0 (#3598 ) (#3841 )	2025-04-29 16:57:22 +08:00
plugin	chore: bump version to 0.19.0 (#3598 ) (#3841 )	2025-04-29 16:57:22 +08:00
quantization	chore: bump version to 0.19.0 (#3598 ) (#3841 )	2025-04-29 16:57:22 +08:00
runtime	feat: Offloading Multimodal embedding table to CPU in Chunked Prefill Mode (#3380 )	2025-04-21 14:31:01 +08:00
scaffolding	feat: fix erros on scaffolding README (#3899 )	2025-04-29 10:15:06 +08:00
serve	feat: Support Top-K logprobs and prompt_logprobs in LLMAPI (#3388 )	2025-05-01 12:47:14 -04:00
tools	test: Fix breaking Phi3 multimodal tests (#3544 )	2025-04-15 08:02:34 +08:00
__init__.py	fix: revert https://github.com/NVIDIA/TensorRT-LLM/pull/3858 (#3928 )	2025-04-29 11:26:13 +08:00
_common.py	Update (#2978 )	2025-03-23 16:39:35 +08:00
_dlpack_utils.py	feat: Add MNNVL MoE A2A support (#3504 )	2025-04-25 17:29:08 +08:00
_ipc_utils.py	fix: Proper error bubbling for PyExecutor (#3321 )	2025-04-15 14:49:46 +08:00
_mnnvl_utils.py	feat: Add MNNVL MoE A2A support (#3504 )	2025-04-25 17:29:08 +08:00
_utils.py	chore: Remove duplicated get_sm_version. (#3935 )	2025-04-30 11:43:53 +08:00
builder.py	chore: remove usernames from comments (#3291 )	2025-04-05 13:44:28 +08:00
disaggregated_params.py	Update TensorRT-LLM (#2936 )	2025-03-18 21:25:19 +08:00
functional.py	Unify two versions of AllReduce custom op (#3032 )	2025-04-22 21:58:42 +08:00
graph_rewriting.py	Update TensorRT-LLM (#2755 )	2025-02-11 03:01:00 +00:00
logger.py	Update TensorRT-LLM (#2873 )	2025-03-11 21:13:42 +08:00
lora_manager.py	add passing E2E LoRA flow (#3788 )	2025-04-23 18:38:06 +03:00
mapping.py	Add smart router for moe (#3641 )	2025-04-23 12:21:59 +08:00
module.py	Update (#2978 )	2025-03-23 16:39:35 +08:00
network.py	chore: remove usernames from comments (#3291 )	2025-04-05 13:44:28 +08:00
parameter.py	Update TensorRT-LLM (#2873 )	2025-03-11 21:13:42 +08:00
profiler.py	test [TRTLLM-4477,TRTLLM-4481]: Accuracy test improvement (Part 3.5): Support GSM8K and GPQA (#3483 )	2025-04-22 07:38:16 +08:00
prompt_adapter_manager.py	Update TensorRT-LLM (#2333 )	2024-10-15 15:28:40 +08:00
python_plugin.py	Update TensorRT-LLM (#2755 )	2025-02-11 03:01:00 +00:00
sampling_params.py	feat: LogitsProcessor in PyTorch backend (#3145 )	2025-05-01 14:15:30 -07:00
top_model_mixin.py	Update TensorRT-LLM (#2053 )	2024-07-30 21:25:01 +08:00
version.py	chore: bump version to 0.20.0rc2 (#3949 )	2025-04-30 11:44:43 +08:00