# SGLang parameter compatibility update - Contract: `0993d90919a8230bf3a9a8a9027db49a4aa89dfc0db2129fec5fbe827b378713` - Diff status: `changed` - Parameters: `488` ## Added - `deepep_v2_mode` - `enable_lean_attention` - `hicache_storage_prefetch_retry_max_attempts` - `hicache_storage_prefetch_retry_poll_interval` - `http2_initial_connection_window_size` - `load_publish_endpoint` ## Removed ## Changed - `bf16_gemm_backend`: `['choices', 'help']` - `disable_priority_preemption`: `['family']` - `dynamic_batch_tokenizer_batch_size`: `['family']` - `dynamic_batch_tokenizer_batch_timeout`: `['family']` - `enable_dynamic_batch_tokenizer`: `['family']` - `enable_prefill_delayer`: `['family']` - `enable_priority_scheduling`: `['family']` - `enable_single_batch_overlap`: `['family']` - `enable_torch_compile`: `['family']` - `enable_torch_compile_debug_mode`: `['family']` - `enable_two_batch_overlap`: `['family']` - `encoder_transfer_backend`: `['default']` - `kv_events_config`: `['help']` - `linear_attn_backend`: `['family']` - `linear_attn_decode_backend`: `['family']` - `linear_attn_prefill_backend`: `['family']` - `linear_attn_verify_backend`: `['family']` - `moe_a2a_backend`: `['choices']` - `moe_runner_backend`: `['choices']` - `prefill_delayer_forward_passes_buckets`: `['family']` - `prefill_delayer_max_delay_ms`: `['family']` - `prefill_delayer_max_delay_passes`: `['family']` - `prefill_delayer_queue_min_ratio`: `['family']` - `prefill_delayer_token_usage_low_watermark`: `['family']` - `prefill_delayer_wait_seconds_buckets`: `['family']` - `priority_scheduling_preemption_threshold`: `['family']` - `sampling_backend`: `['choices', 'default']` - `speculative_moe_a2a_backend`: `['choices']` - `speculative_moe_runner_backend`: `['choices']` - `torch_compile_max_bs`: `['family']` - `triton_attention_num_kv_splits`: `['family']` - `triton_attention_reduce_in_fp32`: `['family']` - `triton_attention_split_tile_size`: `['family']` ## Current parameters without a versioned optimization rule - `abort_on_priority_when_disabled` - `admin_api_key` - `allow_auto_truncate` - `allowed_media_domains` - `api_key` - `asr_max_buffer_seconds` - `asr_max_concurrent_sessions` - `attn_cp_size` - `base_gpu_id` - `batch_notify_size` - `bf16_gemm_backend` - `bucket_e2e_request_latency` - `bucket_inter_token_latency` - `bucket_time_to_first_token` - `c128_page_size` - `chat_template` - `checkpoint_engine_wait_weights_before_ready` - `completion_template` - `config` - `constrained_json_disable_any_whitespace` - `constrained_json_whitespace_pattern` - `context_length` - `cp_strategy` - `cpu_offload_gb` - `crash_dump_folder` - `cuda_graph_backend_decode` - `cuda_graph_config` - `custom_sigquit_handler` - `custom_weight_loader` - `dcp_comm_backend` - `dcp_replicate_q_proj` - `dcp_size` - `debug_cuda_graph` - `debug_tensor_dump_input_file` - `debug_tensor_dump_layers` - `debug_tensor_dump_output_folder` - `decode_log_interval` - `decoupled_spec_bind_endpoint` - `decoupled_spec_connect_endpoints` - `decoupled_spec_rank` - `decoupled_spec_role` - `decrypted_config_file` - `decrypted_draft_config_file` - `deepep_config` - `deepep_dispatcher_output_dtype` - `deepep_mode` - `deepep_v2_mode` - `default_chat_template_kwargs` - `default_priority_value` - `delete_ckpt_after_loading` - `detokenizer_worker_num` - `device` - `disable_attn_tp_gather` - `disable_chunked_prefix_cache` - `disable_cuda_graph_padding` - `disable_decode_cuda_graph` - `disable_flashinfer_autotune` - `disable_flashinfer_cutlass_moe_fp4_allgather` - `disable_hybrid_swa_memory` - `disable_outlines_disk_cache` - `disable_prefill_cuda_graph` - `disable_priority_preemption` - `disable_shared_experts_fusion` - `disable_tokenizer_batch_decode` - `disaggregation_bootstrap_port` - `disaggregation_decode_enable_offload_kvcache` - `disaggregation_decode_enable_radix_cache` - `disaggregation_decode_extra_slots` - `disaggregation_decode_polling_interval` - `disaggregation_decode_retraction_backup` - `disaggregation_ib_device` - `disaggregation_mode` - `disaggregation_transfer_backend` - `dist_init_addr` - `dist_timeout` - `dllm_algorithm` - `dllm_algorithm_config` - `dllm_fdfo` - `download_dir` - `dp_size` - `dsa_paged_mqa_logits_backend` - `dsa_topk_backend` - `dsv4_prefill_backend` - `dtype` - `dwdp_size` - `dynamic_batch_tokenizer_batch_size` - `dynamic_batch_tokenizer_batch_timeout` - `elastic_ep_backend` - `elastic_ep_initial_size` - `elastic_ep_scale_timeout` - `enable_adaptive_dispatch_to_encoder` - `enable_aiter_allreduce_fusion` - `enable_attn_tp_input_scattered` - `enable_broadcast_mm_inputs_process` - `enable_cache_report` - `enable_cp_decode_attn_tp` - `enable_cudagraph_gc` - `enable_custom_logit_processor` - `enable_deepseek_v4_fp4_indexer` - `enable_dense_mlp_attn_tp` - `enable_deterministic_inference` - `enable_dp_attention_local_control_broadcast` - `enable_dp_lm_head` - `enable_draft_weights_cpu_backup` - `enable_dsa_cache_layer_split` - `enable_dynamic_batch_tokenizer` - `enable_dynamic_chunking` - `enable_elastic_expert_backup` - `enable_eplb` - `enable_expert_distribution_metrics` - `enable_flexkv` - `enable_forward_pass_metrics` - `enable_fp32_lm_head` - `enable_fused_moe_sum_all_reduce` - `enable_fused_qk_norm_rope` - `enable_hierarchical_cache` - `enable_hisparse` - `enable_http2` - `enable_int8_mamba_checkpoint` - `enable_layerwise_nvtx_marker` - `enable_lean_attention` - `enable_linear_replayssm` - `enable_lmcache` - `enable_lora` - `enable_lora_overlap_loading` - `enable_mamba_cache_stochastic_rounding` - `enable_memory_saver` - `enable_metrics` - `enable_metrics_for_all_schedulers` - `enable_mfu_metrics` - `enable_mis` - `enable_mm_global_cache` - `enable_multi_layer_eagle` - `enable_multimodal` - `enable_nccl_nvls` - `enable_p2p_check` - `enable_page_major_kv_layout` - `enable_pdmux` - `enable_precise_embedding_interpolation` - `enable_prefill_cp` - `enable_prefill_delayer` - `enable_prefix_mm_cache` - `enable_priority_scheduling` - `enable_profile_cuda_graph` - `enable_quant_communications` - `enable_request_time_stats_logging` - `enable_return_hidden_states` - `enable_return_indexer_topk` - `enable_return_routed_experts` - `enable_scattered_sconv` - `enable_session_radix_cache` - `enable_shared_experts_attn_tp` - `enable_single_batch_overlap` - `enable_ssl_refresh` - `enable_streaming_session` - `enable_strict_thinking` - `enable_symm_mem` - `enable_tf32_matmul` - `enable_tokenizer_batch_encode` - `enable_torch_compile` - `enable_torch_compile_debug_mode` - `enable_torch_symm_mem` - `enable_tp_lm_head_all_to_all` - `enable_trace` - `enable_two_batch_overlap` - `enable_unified_memory` - `enable_w4a4_mxfp4_megamoe` - `enable_waterfill` - `enable_weights_cpu_backup` - `encoder_bootstrap_port` - `encoder_only` - `encoder_register_urls` - `encoder_transfer_backend` - `encoder_urls` - `enforce_disable_flashinfer_allreduce_fusion` - `enforce_shared_experts_fusion` - `engine_info_bootstrap_port` - `ep_dispatch_algorithm` - `ep_join_mode` - `ep_join_rank_offset` - `ep_num_redundant_experts` - `eplb_algorithm` - `eplb_min_rebalancing_utilization_threshold` - `eplb_rebalance_layers_per_chunk` - `eplb_rebalance_num_iterations` - `expert_balancedness_report_mode` - `expert_distribution_recorder_buffer_size` - `expert_distribution_recorder_mode` - `experts_shared_outer_loras` - `export_metrics_to_file` - `export_metrics_to_file_dir` - `extra_metric_labels` - `fastapi_root_path` - `file_storage_path` - `flashinfer_allreduce_fusion_backend` - `flashinfer_autotune_skip_ops` - `flashinfer_mla_disable_ragged` - `flashinfer_mxfp4_moe_precision` - `flexkv_config_file` - `forward_hooks` - `forward_pass_metrics_ipc_name` - `forward_pass_metrics_worker_id` - `fp4_gemm_runner_backend` - `fp8_gemm_runner_backend` - `fuseep_mode` - `gated_launch_port` - `gc_threshold` - `gc_warning_threshold_secs` - `generation_tokens_buckets` - `gpu_id_step` - `grammar_backend` - `grpc_port` - `hf_chat_template_name` - `hicache_host_memory_mode` - `hicache_io_backend` - `hicache_mem_layout` - `hicache_ratio` - `hicache_size` - `hicache_storage_backend` - `hicache_storage_backend_extra_config` - `hicache_storage_prefetch_policy` - `hicache_storage_prefetch_retry_max_attempts` - `hicache_storage_prefetch_retry_poll_interval` - `hicache_write_policy` - `hisparse_config` - `host` - `http2_initial_connection_window_size` - `http2_max_concurrent_streams` - `image_processor_backend` - `init_expert_location` - `int8_mamba_ckpt_size` - `is_embedding` - `json_model_override_args` - `kt_cpuinfer` - `kt_max_deferred_experts_per_token` - `kt_method` - `kt_num_gpu_experts` - `kt_threadpool_count` - `kt_weight_path` - `kv_cache_dtype` - `kv_canary` - `kv_canary_real_data` - `kv_canary_sweep_interval` - `kv_events_config` - `language_model_only` - `language_only` - `limit_mm_data_per_request` - `linear_attn_backend` - `linear_attn_verify_backend` - `linear_replayssm_cache_len` - `lmcache_config_file` - `load_balance_method` - `load_format` - `load_publish_endpoint` - `load_snapshot_publish_interval` - `log_level` - `log_level_http` - `log_requests` - `log_requests_format` - `log_requests_level` - `log_requests_target` - `lora_backend` - `lora_drain_wait_threshold` - `lora_eviction_policy` - `lora_paths` - `lora_strict_loading` - `lora_target_modules` - `lora_use_virtual_experts` - `mamba_backend` - `mamba_cache_philox_rounds` - `mamba_max_states_per_path` - `mamba_track_interval` - `max_ep_size` - `max_loaded_loras` - `max_lora_chunk_size` - `max_lora_rank` - `max_loras_per_batch` - `max_mamba_cache_size` - `max_queued_requests` - `max_running_requests` - `max_total_tokens` - `media_url_max_file_size_mb` - `min_free_slots_delay` - `mlx_enable_sampling` - `mm_attention_backend` - `mm_enable_dp_encoder` - `mm_feature_transport` - `mm_global_cache_backend` - `mm_io_worker_num` - `mm_preprocess_cache_size_mb` - `mm_process_config` - `mm_processor_worker_num` - `model_checksum` - `model_config_parser` - `model_impl` - `model_loader_extra_config` - `model_path` - `modelexpress_config` - `modelopt_checkpoint_restore_path` - `modelopt_checkpoint_save_path` - `modelopt_export_path` - `modelopt_quant` - `moe_a2a_backend` - `moe_dense_tp_size` - `moe_dp_size` - `mooncake_ib_device` - `msprobe_dump_config` - `nccl_port` - `nnodes` - `node_rank` - `num_reserved_decode_tokens` - `numa_node` - `offload_group_size` - `offload_mode` - `offload_num_in_group` - `offload_prefetch_step` - `optimistic_prefill_attempts` - `otlp_traces_endpoint` - `pdmux_config_path` - `port` - `pp_async_batch_depth` - `pp_max_micro_batch_size` - `pp_size` - `pre_warm_nccl` - `preferred_sampling_params` - `prefill_decode_interval` - `prefill_delayer_forward_passes_buckets` - `prefill_delayer_max_delay_ms` - `prefill_delayer_max_delay_passes` - `prefill_delayer_queue_min_ratio` - `prefill_delayer_token_usage_low_watermark` - `prefill_delayer_wait_seconds_buckets` - `prefill_max_requests` - `prefill_only_disable_kv_cache` - `priority_scheduling_preemption_threshold` - `prompt_tokens_buckets` - `quantization` - `quantization_param_path` - `quantize_and_serve` - `radix_cache_backend` - `radix_eviction_policy` - `random_seed` - `reasoning_parser` - `remote_instance_weight_loader_backend` - `remote_instance_weight_loader_seed_instance_ip` - `remote_instance_weight_loader_seed_instance_service_port` - `remote_instance_weight_loader_send_weights_group_ports` - `remote_instance_weight_loader_start_seed_via_transfer_engine` - `retraction_policy` - `return_hidden_states_mode` - `revision` - `rl_on_policy_target` - `rl_quant_profile` - `sampling_backend` - `sampling_defaults` - `schedule_low_priority_values_first` - `served_model_name` - `show_time_cost` - `sidecar` - `sidecar_args` - `skip_server_warmup` - `skip_tokenizer_init` - `sleep_on_idle` - `sm_group_num` - `smg_http_sidecar_port` - `soft_watchdog_timeout` - `spec_trace_dir` - `speculative_accept_threshold_acc` - `speculative_accept_threshold_single` - `speculative_adaptive` - `speculative_adaptive_config` - `speculative_attention_mode` - `speculative_dflash_block_size` - `speculative_draft_attention_backend` - `speculative_draft_kv_cache_dtype` - `speculative_draft_load_format` - `speculative_draft_model_path` - `speculative_draft_model_quantization` - `speculative_draft_model_revision` - `speculative_draft_window_size` - `speculative_dsa_topk_backend` - `speculative_dspark_align_verify_tokens_to_graph_tier` - `speculative_dspark_block_size` - `speculative_dspark_confidence_sts_path` - `speculative_dspark_sps_table_path` - `speculative_eagle_topk` - `speculative_moe_a2a_backend` - `speculative_moe_runner_backend` - `speculative_ngram_capacity` - `speculative_ngram_external_corpus_max_tokens` - `speculative_ngram_external_corpus_path` - `speculative_ngram_external_sam_budget` - `speculative_ngram_match_type` - `speculative_ngram_max_bfs_breadth` - `speculative_ngram_max_trie_depth` - `speculative_ngram_min_bfs_breadth` - `speculative_skip_dp_mlp_sync` - `speculative_token_map` - `speculative_use_rejection_sampling` - `ssl_ca_certs` - `ssl_certfile` - `ssl_keyfile` - `ssl_keyfile_password` - `startup_weight_load_mode` - `stat_loggers` - `stream_interval` - `stream_response_default_include_usage` - `strip_thinking_cache` - `swa_full_tokens_ratio` - `tbo_token_distribution_threshold` - `tokenizer_backend` - `tokenizer_metrics_allowed_custom_labels` - `tokenizer_metrics_custom_labels_header` - `tokenizer_mode` - `tokenizer_path` - `tokenizer_worker_num` - `tool_call_parser` - `tool_server` - `torch_compile_max_bs` - `trace_modules` - `triton_attention_num_kv_splits` - `triton_attention_reduce_in_fp32` - `triton_attention_split_tile_size` - `trust_mm_content_hashes` - `trust_remote_code` - `use_ray` - `uvicorn_access_log_exclude_prefixes` - `warmups` - `watchdog_timeout` - `weight_cache_mode` - `weight_cache_socket` - `weight_cache_timeout` - `weight_loader_disable_mmap` - `weight_loader_drop_cache_after_load` - `weight_loader_prefetch_checkpoints` - `weight_loader_prefetch_num_threads` - `weight_version` Unknown parameters are not automatically trusted. Run InferOpt semantic/safety analysis and real benchmarks before promoting them to validated rules.
SGLang parameter compatibility update
0993d90919a8230bf3a9a8a9027db49a4aa89dfc0db2129fec5fbe827b378713changed488Added
deepep_v2_modeenable_lean_attentionhicache_storage_prefetch_retry_max_attemptshicache_storage_prefetch_retry_poll_intervalhttp2_initial_connection_window_sizeload_publish_endpointRemoved
Changed
bf16_gemm_backend:['choices', 'help']disable_priority_preemption:['family']dynamic_batch_tokenizer_batch_size:['family']dynamic_batch_tokenizer_batch_timeout:['family']enable_dynamic_batch_tokenizer:['family']enable_prefill_delayer:['family']enable_priority_scheduling:['family']enable_single_batch_overlap:['family']enable_torch_compile:['family']enable_torch_compile_debug_mode:['family']enable_two_batch_overlap:['family']encoder_transfer_backend:['default']kv_events_config:['help']linear_attn_backend:['family']linear_attn_decode_backend:['family']linear_attn_prefill_backend:['family']linear_attn_verify_backend:['family']moe_a2a_backend:['choices']moe_runner_backend:['choices']prefill_delayer_forward_passes_buckets:['family']prefill_delayer_max_delay_ms:['family']prefill_delayer_max_delay_passes:['family']prefill_delayer_queue_min_ratio:['family']prefill_delayer_token_usage_low_watermark:['family']prefill_delayer_wait_seconds_buckets:['family']priority_scheduling_preemption_threshold:['family']sampling_backend:['choices', 'default']speculative_moe_a2a_backend:['choices']speculative_moe_runner_backend:['choices']torch_compile_max_bs:['family']triton_attention_num_kv_splits:['family']triton_attention_reduce_in_fp32:['family']triton_attention_split_tile_size:['family']Current parameters without a versioned optimization rule
abort_on_priority_when_disabledadmin_api_keyallow_auto_truncateallowed_media_domainsapi_keyasr_max_buffer_secondsasr_max_concurrent_sessionsattn_cp_sizebase_gpu_idbatch_notify_sizebf16_gemm_backendbucket_e2e_request_latencybucket_inter_token_latencybucket_time_to_first_tokenc128_page_sizechat_templatecheckpoint_engine_wait_weights_before_readycompletion_templateconfigconstrained_json_disable_any_whitespaceconstrained_json_whitespace_patterncontext_lengthcp_strategycpu_offload_gbcrash_dump_foldercuda_graph_backend_decodecuda_graph_configcustom_sigquit_handlercustom_weight_loaderdcp_comm_backenddcp_replicate_q_projdcp_sizedebug_cuda_graphdebug_tensor_dump_input_filedebug_tensor_dump_layersdebug_tensor_dump_output_folderdecode_log_intervaldecoupled_spec_bind_endpointdecoupled_spec_connect_endpointsdecoupled_spec_rankdecoupled_spec_roledecrypted_config_filedecrypted_draft_config_filedeepep_configdeepep_dispatcher_output_dtypedeepep_modedeepep_v2_modedefault_chat_template_kwargsdefault_priority_valuedelete_ckpt_after_loadingdetokenizer_worker_numdevicedisable_attn_tp_gatherdisable_chunked_prefix_cachedisable_cuda_graph_paddingdisable_decode_cuda_graphdisable_flashinfer_autotunedisable_flashinfer_cutlass_moe_fp4_allgatherdisable_hybrid_swa_memorydisable_outlines_disk_cachedisable_prefill_cuda_graphdisable_priority_preemptiondisable_shared_experts_fusiondisable_tokenizer_batch_decodedisaggregation_bootstrap_portdisaggregation_decode_enable_offload_kvcachedisaggregation_decode_enable_radix_cachedisaggregation_decode_extra_slotsdisaggregation_decode_polling_intervaldisaggregation_decode_retraction_backupdisaggregation_ib_devicedisaggregation_modedisaggregation_transfer_backenddist_init_addrdist_timeoutdllm_algorithmdllm_algorithm_configdllm_fdfodownload_dirdp_sizedsa_paged_mqa_logits_backenddsa_topk_backenddsv4_prefill_backenddtypedwdp_sizedynamic_batch_tokenizer_batch_sizedynamic_batch_tokenizer_batch_timeoutelastic_ep_backendelastic_ep_initial_sizeelastic_ep_scale_timeoutenable_adaptive_dispatch_to_encoderenable_aiter_allreduce_fusionenable_attn_tp_input_scatteredenable_broadcast_mm_inputs_processenable_cache_reportenable_cp_decode_attn_tpenable_cudagraph_gcenable_custom_logit_processorenable_deepseek_v4_fp4_indexerenable_dense_mlp_attn_tpenable_deterministic_inferenceenable_dp_attention_local_control_broadcastenable_dp_lm_headenable_draft_weights_cpu_backupenable_dsa_cache_layer_splitenable_dynamic_batch_tokenizerenable_dynamic_chunkingenable_elastic_expert_backupenable_eplbenable_expert_distribution_metricsenable_flexkvenable_forward_pass_metricsenable_fp32_lm_headenable_fused_moe_sum_all_reduceenable_fused_qk_norm_ropeenable_hierarchical_cacheenable_hisparseenable_http2enable_int8_mamba_checkpointenable_layerwise_nvtx_markerenable_lean_attentionenable_linear_replayssmenable_lmcacheenable_loraenable_lora_overlap_loadingenable_mamba_cache_stochastic_roundingenable_memory_saverenable_metricsenable_metrics_for_all_schedulersenable_mfu_metricsenable_misenable_mm_global_cacheenable_multi_layer_eagleenable_multimodalenable_nccl_nvlsenable_p2p_checkenable_page_major_kv_layoutenable_pdmuxenable_precise_embedding_interpolationenable_prefill_cpenable_prefill_delayerenable_prefix_mm_cacheenable_priority_schedulingenable_profile_cuda_graphenable_quant_communicationsenable_request_time_stats_loggingenable_return_hidden_statesenable_return_indexer_topkenable_return_routed_expertsenable_scattered_sconvenable_session_radix_cacheenable_shared_experts_attn_tpenable_single_batch_overlapenable_ssl_refreshenable_streaming_sessionenable_strict_thinkingenable_symm_memenable_tf32_matmulenable_tokenizer_batch_encodeenable_torch_compileenable_torch_compile_debug_modeenable_torch_symm_memenable_tp_lm_head_all_to_allenable_traceenable_two_batch_overlapenable_unified_memoryenable_w4a4_mxfp4_megamoeenable_waterfillenable_weights_cpu_backupencoder_bootstrap_portencoder_onlyencoder_register_urlsencoder_transfer_backendencoder_urlsenforce_disable_flashinfer_allreduce_fusionenforce_shared_experts_fusionengine_info_bootstrap_portep_dispatch_algorithmep_join_modeep_join_rank_offsetep_num_redundant_expertseplb_algorithmeplb_min_rebalancing_utilization_thresholdeplb_rebalance_layers_per_chunkeplb_rebalance_num_iterationsexpert_balancedness_report_modeexpert_distribution_recorder_buffer_sizeexpert_distribution_recorder_modeexperts_shared_outer_lorasexport_metrics_to_fileexport_metrics_to_file_dirextra_metric_labelsfastapi_root_pathfile_storage_pathflashinfer_allreduce_fusion_backendflashinfer_autotune_skip_opsflashinfer_mla_disable_raggedflashinfer_mxfp4_moe_precisionflexkv_config_fileforward_hooksforward_pass_metrics_ipc_nameforward_pass_metrics_worker_idfp4_gemm_runner_backendfp8_gemm_runner_backendfuseep_modegated_launch_portgc_thresholdgc_warning_threshold_secsgeneration_tokens_bucketsgpu_id_stepgrammar_backendgrpc_porthf_chat_template_namehicache_host_memory_modehicache_io_backendhicache_mem_layouthicache_ratiohicache_sizehicache_storage_backendhicache_storage_backend_extra_confighicache_storage_prefetch_policyhicache_storage_prefetch_retry_max_attemptshicache_storage_prefetch_retry_poll_intervalhicache_write_policyhisparse_confighosthttp2_initial_connection_window_sizehttp2_max_concurrent_streamsimage_processor_backendinit_expert_locationint8_mamba_ckpt_sizeis_embeddingjson_model_override_argskt_cpuinferkt_max_deferred_experts_per_tokenkt_methodkt_num_gpu_expertskt_threadpool_countkt_weight_pathkv_cache_dtypekv_canarykv_canary_real_datakv_canary_sweep_intervalkv_events_configlanguage_model_onlylanguage_onlylimit_mm_data_per_requestlinear_attn_backendlinear_attn_verify_backendlinear_replayssm_cache_lenlmcache_config_fileload_balance_methodload_formatload_publish_endpointload_snapshot_publish_intervallog_levellog_level_httplog_requestslog_requests_formatlog_requests_levellog_requests_targetlora_backendlora_drain_wait_thresholdlora_eviction_policylora_pathslora_strict_loadinglora_target_moduleslora_use_virtual_expertsmamba_backendmamba_cache_philox_roundsmamba_max_states_per_pathmamba_track_intervalmax_ep_sizemax_loaded_lorasmax_lora_chunk_sizemax_lora_rankmax_loras_per_batchmax_mamba_cache_sizemax_queued_requestsmax_running_requestsmax_total_tokensmedia_url_max_file_size_mbmin_free_slots_delaymlx_enable_samplingmm_attention_backendmm_enable_dp_encodermm_feature_transportmm_global_cache_backendmm_io_worker_nummm_preprocess_cache_size_mbmm_process_configmm_processor_worker_nummodel_checksummodel_config_parsermodel_implmodel_loader_extra_configmodel_pathmodelexpress_configmodelopt_checkpoint_restore_pathmodelopt_checkpoint_save_pathmodelopt_export_pathmodelopt_quantmoe_a2a_backendmoe_dense_tp_sizemoe_dp_sizemooncake_ib_devicemsprobe_dump_confignccl_portnnodesnode_ranknum_reserved_decode_tokensnuma_nodeoffload_group_sizeoffload_modeoffload_num_in_groupoffload_prefetch_stepoptimistic_prefill_attemptsotlp_traces_endpointpdmux_config_pathportpp_async_batch_depthpp_max_micro_batch_sizepp_sizepre_warm_ncclpreferred_sampling_paramsprefill_decode_intervalprefill_delayer_forward_passes_bucketsprefill_delayer_max_delay_msprefill_delayer_max_delay_passesprefill_delayer_queue_min_ratioprefill_delayer_token_usage_low_watermarkprefill_delayer_wait_seconds_bucketsprefill_max_requestsprefill_only_disable_kv_cachepriority_scheduling_preemption_thresholdprompt_tokens_bucketsquantizationquantization_param_pathquantize_and_serveradix_cache_backendradix_eviction_policyrandom_seedreasoning_parserremote_instance_weight_loader_backendremote_instance_weight_loader_seed_instance_ipremote_instance_weight_loader_seed_instance_service_portremote_instance_weight_loader_send_weights_group_portsremote_instance_weight_loader_start_seed_via_transfer_engineretraction_policyreturn_hidden_states_moderevisionrl_on_policy_targetrl_quant_profilesampling_backendsampling_defaultsschedule_low_priority_values_firstserved_model_nameshow_time_costsidecarsidecar_argsskip_server_warmupskip_tokenizer_initsleep_on_idlesm_group_numsmg_http_sidecar_portsoft_watchdog_timeoutspec_trace_dirspeculative_accept_threshold_accspeculative_accept_threshold_singlespeculative_adaptivespeculative_adaptive_configspeculative_attention_modespeculative_dflash_block_sizespeculative_draft_attention_backendspeculative_draft_kv_cache_dtypespeculative_draft_load_formatspeculative_draft_model_pathspeculative_draft_model_quantizationspeculative_draft_model_revisionspeculative_draft_window_sizespeculative_dsa_topk_backendspeculative_dspark_align_verify_tokens_to_graph_tierspeculative_dspark_block_sizespeculative_dspark_confidence_sts_pathspeculative_dspark_sps_table_pathspeculative_eagle_topkspeculative_moe_a2a_backendspeculative_moe_runner_backendspeculative_ngram_capacityspeculative_ngram_external_corpus_max_tokensspeculative_ngram_external_corpus_pathspeculative_ngram_external_sam_budgetspeculative_ngram_match_typespeculative_ngram_max_bfs_breadthspeculative_ngram_max_trie_depthspeculative_ngram_min_bfs_breadthspeculative_skip_dp_mlp_syncspeculative_token_mapspeculative_use_rejection_samplingssl_ca_certsssl_certfilessl_keyfilessl_keyfile_passwordstartup_weight_load_modestat_loggersstream_intervalstream_response_default_include_usagestrip_thinking_cacheswa_full_tokens_ratiotbo_token_distribution_thresholdtokenizer_backendtokenizer_metrics_allowed_custom_labelstokenizer_metrics_custom_labels_headertokenizer_modetokenizer_pathtokenizer_worker_numtool_call_parsertool_servertorch_compile_max_bstrace_modulestriton_attention_num_kv_splitstriton_attention_reduce_in_fp32triton_attention_split_tile_sizetrust_mm_content_hashestrust_remote_codeuse_rayuvicorn_access_log_exclude_prefixeswarmupswatchdog_timeoutweight_cache_modeweight_cache_socketweight_cache_timeoutweight_loader_disable_mmapweight_loader_drop_cache_after_loadweight_loader_prefetch_checkpointsweight_loader_prefetch_num_threadsweight_versionUnknown parameters are not automatically trusted. Run InferOpt semantic/safety analysis and real benchmarks before promoting them to validated rules.