TensorRT-LLMs

kanshan/TensorRT-LLMs

Fork 0

mirror of https://github.com/NVIDIA/TensorRT-LLM.git synced 2026-01-26 21:53:30 +08:00

Commit Graph

Author	SHA1	Message	Date
Yukun He	bed5bc9f2e	[None][chore] Wrap the swiglu into custom op to avoid redundant device copy. (#7021 ) A redundant D2D copy is observed when enabling torch.compile for the Llama model due to the swiglu triton kernel, which brings perf overhead. Use a custom op to wrap the swiglu op to avoid this overhead. Signed-off-by: Yukun He <23156053+hyukn@users.noreply.github.com>	2025-08-27 13:02:10 +08:00
JunyiXu-nv	13e0214fe0	[TRTLLM-6263][feat] Enable fp8 SwiGLU to minimize host overhead (#6540 ) Signed-off-by: Junyi Xu <junyix@nvidia.com>	2025-08-06 10:42:19 +08:00

Author

SHA1

Message

Date

Yukun He

bed5bc9f2e

[None][chore] Wrap the swiglu into custom op to avoid redundant device copy. (#7021 )

A redundant D2D copy is observed when enabling torch.compile for the Llama model due to the swiglu triton kernel, which brings perf overhead. Use a custom op to wrap the swiglu op to avoid this overhead.

Signed-off-by: Yukun He <23156053+hyukn@users.noreply.github.com>

2025-08-27 13:02:10 +08:00

JunyiXu-nv

13e0214fe0

[TRTLLM-6263][feat] Enable fp8 SwiGLU to minimize host overhead (#6540 )

Signed-off-by: Junyi Xu <junyix@nvidia.com>

2025-08-06 10:42:19 +08:00

2 Commits