diff --git a/docs/features/tool_calling.md b/docs/features/tool_calling.md index 43010c406f5..1d10a94c712 100644 --- a/docs/features/tool_calling.md +++ b/docs/features/tool_calling.md @@ -109,18 +109,18 @@ vLLM supports the `tool_choice='none'` option in the chat completion API. When t ## Constrained Decoding Behavior -Whether vLLM enforces the tool parameter schema during generation depends on the `tool_choice` mode: +Whether vLLM enforces the tool parameter schema during generation depends on the `tool_choice` mode and the per-tool `strict` field: | `tool_choice` value | Schema-constrained decoding | Behavior | | --- | --- | --- | | Named function | Yes (via structured outputs backend) | Arguments are guaranteed to be valid JSON conforming to the function's parameter schema. | | `"required"` | Yes (via structured outputs backend) | Same as named function. The model must produce at least one tool call. | -| `"auto"` | Depends on the parser | Model-specific structural-tag parsers can constrain tool-call arguments with structured outputs. Other parsers generate freely and extract tool calls from raw text. | +| `"auto"` | Only when `strict: true` is set on at least one tool | Structural-tag parsers constrain tool-call arguments when a tool opts in with `strict: true`. Without it, the model generates freely and tool calls are extracted from raw text. | | `"none"` | N/A | No tool calls are produced. | ### Strict Mode -Strict tool calling makes function-call arguments adhere to the function schema instead of relying only on best-effort parsing. vLLM implements strict tool calling for structural-tag based tool parsers by using the structured outputs backend under the hood. +For `tool_choice="required"` or named function calling, structural-tag constraints are always applied regardless of the `strict` field. For `tool_choice="auto"`, setting `strict: true` on at least one tool opts in to structural-tag constraints; without it, the model generates freely and tool calls are extracted from raw text. The `strict` field is supported across all three API surfaces: Chat Completion, Responses, and Anthropic Messages. For best compatibility with strict schema enforcement, define tool parameter schemas in the OpenAI strict-schema style: @@ -128,16 +128,12 @@ For best compatibility with strict schema enforcement, define tool parameter sch * Mark all fields in `properties` as required. * Represent optional fields by allowing `null`, for example `{"type": ["string", "null"]}`. -vLLM controls structural-tag strict tool calling with the `VLLM_ENFORCE_STRICT_TOOL_CALLING` environment variable. It defaults to `true`. +vLLM also provides a global toggle via the `VLLM_ENFORCE_STRICT_TOOL_CALLING` environment variable (defaults to `true`). When set to `false`, vLLM does not attach structural tags for tool calling regardless of the per-tool `strict` field. This environment variable only affects structural-tag based tool calling; it does not change schema-derived structured outputs used by named function calling or `tool_choice="required"`. ```bash VLLM_ENFORCE_STRICT_TOOL_CALLING=false vllm serve ... ``` -When this variable is `true`, structural-tag based tool parsers attach a structural tag to the request, so the structured outputs backend can constrain the model-specific tool-call format and function-call arguments. When it is `false`, vLLM does not attach structural tags for tool calling. In that case, `tool_choice="auto"` falls back to best-effort parser extraction from the raw model output, and no structural-tag constraint is applied. - -This environment variable only affects structural-tag based tool calling. It does not change schema-derived structured outputs used by named function calling or `tool_choice="required"`. - ## Automatic Function Calling To enable this feature, you should set the following flags: @@ -156,7 +152,7 @@ from HuggingFace; and you can find an example of this in a `tokenizer_config.jso If your favorite tool-calling model is not supported, please feel free to contribute a parser & tool use chat template! !!! note - With `tool_choice="auto"`, schema-level constraint depends on the selected parser and `VLLM_ENFORCE_STRICT_TOOL_CALLING`. Structural-tag parsers can enforce tool-call constraints when it is `true`; when it is `false`, or when the selected parser has no structural-tag support, vLLM extracts tool calls from raw text, so arguments may occasionally be malformed or violate the function's parameter schema. + With `tool_choice="auto"`, schema-level constraint requires both `VLLM_ENFORCE_STRICT_TOOL_CALLING=true` (the default) and at least one tool with `strict: true`. When these conditions are met and the selected parser supports structural tags, vLLM constrains tool-call arguments. Otherwise, vLLM extracts tool calls from raw text, so arguments may occasionally be malformed or violate the function's parameter schema. ### Hermes Models (`hermes`) diff --git a/tests/tool_parsers/test_deepseekv4_tool_parser.py b/tests/tool_parsers/test_deepseekv4_tool_parser.py index ab66d6e64cd..80e3357b68b 100644 --- a/tests/tool_parsers/test_deepseekv4_tool_parser.py +++ b/tests/tool_parsers/test_deepseekv4_tool_parser.py @@ -216,14 +216,32 @@ def test_streaming_emits_incremental_argument_chunks(): } +def _with_strict( + tools: list[ChatCompletionToolsParam], +) -> list[ChatCompletionToolsParam]: + return [ + ChatCompletionToolsParam( + type=t.type, + function=FunctionDefinition( + name=t.function.name, + description=t.function.description, + parameters=t.function.parameters, + strict=True, + ), + ) + for t in tools + ] + + def test_get_vllm_registry_structural_tag_returns_structural_tag( sample_tools: list[ChatCompletionToolsParam], ) -> None: parser = make_parser() + strict_tools = _with_strict(sample_tools) req = ChatCompletionRequest( messages=[], model="m", - tools=sample_tools, + tools=strict_tools, tool_choice="auto", ) tag = parser.get_structural_tag(req) diff --git a/tests/tool_parsers/test_qwen3coder_tool_parser.py b/tests/tool_parsers/test_qwen3coder_tool_parser.py index 90c5013431e..ac770ff8e5b 100644 --- a/tests/tool_parsers/test_qwen3coder_tool_parser.py +++ b/tests/tool_parsers/test_qwen3coder_tool_parser.py @@ -14,6 +14,7 @@ from vllm.entrypoints.openai.chat_completion.protocol import ( ChatCompletionNamedToolChoiceParam, ChatCompletionRequest, ChatCompletionToolsParam, + FunctionDefinition, ) from vllm.entrypoints.openai.engine.protocol import ( DeltaMessage, @@ -115,6 +116,23 @@ def sample_tools(request): ] +def _with_strict( + tools: list[ChatCompletionToolsParam], +) -> list[ChatCompletionToolsParam]: + return [ + ChatCompletionToolsParam( + type=t.type, + function=FunctionDefinition( + name=t.function.name, + description=t.function.description, + parameters=t.function.parameters, + strict=True, + ), + ) + for t in tools + ] + + def _as_chat_completion_tools( tools: list[ChatCompletionToolsParam | FunctionTool], ) -> list[ChatCompletionToolsParam]: @@ -1323,10 +1341,11 @@ def test_get_vllm_registry_structural_tag_returns_structural_tag( sample_tools: list[ChatCompletionToolsParam], ) -> None: request_tools = _as_chat_completion_tools(sample_tools) + strict_tools = _with_strict(request_tools) req = ChatCompletionRequest( messages=[], model="m", - tools=request_tools, + tools=strict_tools, tool_choice="auto", ) tag = qwen3_tool_parser.get_structural_tag(req) @@ -1364,10 +1383,11 @@ def test_adjust_request_auto_uses_vllm_registry_structural_tag( tool_parser_cls = Qwen3EngineToolParser request_tools = _as_chat_completion_tools(sample_tools) + strict_tools = _with_strict(request_tools) req = ChatCompletionRequest( messages=[], model="m", - tools=request_tools, + tools=strict_tools, tool_choice="auto", include_reasoning=include_reasoning, ) diff --git a/tests/tool_parsers/test_structural_tag_registry.py b/tests/tool_parsers/test_structural_tag_registry.py index 530a812566c..bd84b2cbbfa 100644 --- a/tests/tool_parsers/test_structural_tag_registry.py +++ b/tests/tool_parsers/test_structural_tag_registry.py @@ -51,6 +51,24 @@ def sample_tools() -> list[ChatCompletionToolsParam]: ] +@pytest.fixture +def sample_tools_strict() -> list[ChatCompletionToolsParam]: + return [ + ChatCompletionToolsParam( + type="function", + function={ + "name": "get_weather", + "strict": True, + "parameters": { + "type": "object", + "properties": {"city": {"type": "string"}}, + "required": ["city"], + }, + }, + ) + ] + + def test_supported_structural_tag_models_include_vllm_builtins(): assert SUPPORTED_STRUCTURAL_TAG_MODELS == ( XGRAMMAR_BUILTIN_STRUCTURAL_TAG_MODELS | VLLM_BUILTIN_STRUCTURAL_TAG_MODELS @@ -61,11 +79,11 @@ def test_supported_structural_tag_models_include_vllm_builtins(): @pytest.mark.parametrize("model", sorted(XGRAMMAR_BUILTIN_STRUCTURAL_TAG_MODELS)) def test_get_model_structural_tag_supports_all_xgrammar_builtins( model: str, - sample_tools: list[ChatCompletionToolsParam], + sample_tools_strict: list[ChatCompletionToolsParam], ): tag = get_model_structural_tag( model=model, - tools=sample_tools, + tools=sample_tools_strict, tool_choice="auto", reasoning=False, ) @@ -219,7 +237,7 @@ def test_non_structural_tag_parser_uses_schema_constraints( def test_get_structural_tag_disables_reasoning( monkeypatch: pytest.MonkeyPatch, - sample_tools: list[ChatCompletionToolsParam], + sample_tools_strict: list[ChatCompletionToolsParam], ): captured: list[bool] = [] @@ -235,10 +253,10 @@ def test_get_structural_tag_disables_reasoning( request = ChatCompletionRequest( messages=[], model="m", - tools=sample_tools, + tools=sample_tools_strict, tool_choice="auto", ) - parser = Qwen3EngineToolParser(MagicMock(), tools=sample_tools) + parser = Qwen3EngineToolParser(MagicMock(), tools=sample_tools_strict) parser.get_structural_tag(request) @@ -247,7 +265,7 @@ def test_get_structural_tag_disables_reasoning( def test_unified_parser_get_structural_tag_disables_reasoning( monkeypatch: pytest.MonkeyPatch, - sample_tools: list[ChatCompletionToolsParam], + sample_tools_strict: list[ChatCompletionToolsParam], ): captured: list[bool] = [] @@ -266,10 +284,10 @@ def test_unified_parser_get_structural_tag_disables_reasoning( request = ChatCompletionRequest( messages=[], model="m", - tools=sample_tools, + tools=sample_tools_strict, tool_choice="auto", ) - parser = TestParser(MagicMock(), tools=sample_tools) + parser = TestParser(MagicMock(), tools=sample_tools_strict) parser.reasoning_parser = MagicMock(adjust_request=lambda request: request) parser.adjust_request(request) @@ -279,7 +297,7 @@ def test_unified_parser_get_structural_tag_disables_reasoning( def test_xgrammar_function_parameters_are_preserved( monkeypatch: pytest.MonkeyPatch, - sample_tools: list[ChatCompletionToolsParam], + sample_tools_strict: list[ChatCompletionToolsParam], ): captured: list[list[dict]] = [] @@ -294,15 +312,31 @@ def test_xgrammar_function_parameters_are_preserved( get_model_structural_tag( model="llama", - tools=sample_tools, + tools=sample_tools_strict, tool_choice="auto", reasoning=False, ) assert ( - captured[0][0]["function"]["parameters"] == sample_tools[0].function.parameters + captured[0][0]["function"]["parameters"] + == sample_tools_strict[0].function.parameters ) - assert sample_tools[0].function.parameters is not None + assert sample_tools_strict[0].function.parameters is not None + + +@pytest.mark.parametrize("model", sorted(XGRAMMAR_BUILTIN_STRUCTURAL_TAG_MODELS)) +def test_auto_tool_choice_skips_structural_tag_without_strict( + model: str, + sample_tools: list[ChatCompletionToolsParam], +): + tag = get_model_structural_tag( + model=model, + tools=sample_tools, + tool_choice="auto", + reasoning=False, + ) + + assert tag is None def test_get_function_parameters_relaxes_function_strict_false(): diff --git a/vllm/entrypoints/anthropic/protocol.py b/vllm/entrypoints/anthropic/protocol.py index 279f3625345..ae0dd08660d 100644 --- a/vllm/entrypoints/anthropic/protocol.py +++ b/vllm/entrypoints/anthropic/protocol.py @@ -75,6 +75,7 @@ class AnthropicTool(BaseModel): name: str description: str | None = None input_schema: dict[str, Any] + strict: bool | None = None defer_loading: bool | None = None @field_validator("input_schema") diff --git a/vllm/entrypoints/anthropic/serving.py b/vllm/entrypoints/anthropic/serving.py index 3dce10695b5..229b7acda62 100644 --- a/vllm/entrypoints/anthropic/serving.py +++ b/vllm/entrypoints/anthropic/serving.py @@ -462,6 +462,7 @@ class AnthropicServingMessages(OpenAIServingChat): "name": tool.name, "description": tool.description, "parameters": tool.input_schema, + "strict": tool.strict, "defer_loading": tool.defer_loading, }, } diff --git a/vllm/entrypoints/openai/engine/protocol.py b/vllm/entrypoints/openai/engine/protocol.py index 3cd998780f9..d86c77561db 100644 --- a/vllm/entrypoints/openai/engine/protocol.py +++ b/vllm/entrypoints/openai/engine/protocol.py @@ -247,11 +247,14 @@ class FunctionDefinition(OpenAIBaseModel): name: str description: str | None = None parameters: dict[str, Any] | None = None + strict: bool | None = None defer_loading: bool | None = None @model_serializer(mode="wrap") def _serialize(self, handler): data = handler(self) + if self.strict is None: + data.pop("strict", None) if self.defer_loading is None: data.pop("defer_loading", None) return data diff --git a/vllm/tool_parsers/structural_tag_registry.py b/vllm/tool_parsers/structural_tag_registry.py index 13491e95dfc..99c92f8f0a2 100644 --- a/vllm/tool_parsers/structural_tag_registry.py +++ b/vllm/tool_parsers/structural_tag_registry.py @@ -84,6 +84,17 @@ def register_vllm_structural_tag(model: str): return decorator +def _any_tool_strict( + tools: Sequence[ChatCompletionToolsParam | ResponsesTool], +) -> bool: + for tool in tools: + if isinstance(tool, FunctionTool) and tool.strict is True: + return True + if isinstance(tool, ChatCompletionToolsParam) and tool.function.strict is True: + return True + return False + + def get_model_structural_tag( model: str, tools: Sequence[ChatCompletionToolsParam | ResponsesTool] | None, @@ -95,6 +106,9 @@ def get_model_structural_tag( if not tools or tool_choice == "none": return None + if tool_choice == "auto" and not _any_tool_strict(tools): + return None + dumped_tools = [_dump_tool_for_xgrammar(tool) for tool in tools] dumped_tool_choice = _dump_tool_choice_for_xgrammar(tool_choice)