Skip to content

SGLang parameter compatibility update #2

Description

@github-actions

SGLang parameter compatibility update

  • Contract: 0993d90919a8230bf3a9a8a9027db49a4aa89dfc0db2129fec5fbe827b378713
  • Diff status: changed
  • Parameters: 488

Added

  • deepep_v2_mode
  • enable_lean_attention
  • hicache_storage_prefetch_retry_max_attempts
  • hicache_storage_prefetch_retry_poll_interval
  • http2_initial_connection_window_size
  • load_publish_endpoint

Removed

Changed

  • bf16_gemm_backend: ['choices', 'help']
  • disable_priority_preemption: ['family']
  • dynamic_batch_tokenizer_batch_size: ['family']
  • dynamic_batch_tokenizer_batch_timeout: ['family']
  • enable_dynamic_batch_tokenizer: ['family']
  • enable_prefill_delayer: ['family']
  • enable_priority_scheduling: ['family']
  • enable_single_batch_overlap: ['family']
  • enable_torch_compile: ['family']
  • enable_torch_compile_debug_mode: ['family']
  • enable_two_batch_overlap: ['family']
  • encoder_transfer_backend: ['default']
  • kv_events_config: ['help']
  • linear_attn_backend: ['family']
  • linear_attn_decode_backend: ['family']
  • linear_attn_prefill_backend: ['family']
  • linear_attn_verify_backend: ['family']
  • moe_a2a_backend: ['choices']
  • moe_runner_backend: ['choices']
  • prefill_delayer_forward_passes_buckets: ['family']
  • prefill_delayer_max_delay_ms: ['family']
  • prefill_delayer_max_delay_passes: ['family']
  • prefill_delayer_queue_min_ratio: ['family']
  • prefill_delayer_token_usage_low_watermark: ['family']
  • prefill_delayer_wait_seconds_buckets: ['family']
  • priority_scheduling_preemption_threshold: ['family']
  • sampling_backend: ['choices', 'default']
  • speculative_moe_a2a_backend: ['choices']
  • speculative_moe_runner_backend: ['choices']
  • torch_compile_max_bs: ['family']
  • triton_attention_num_kv_splits: ['family']
  • triton_attention_reduce_in_fp32: ['family']
  • triton_attention_split_tile_size: ['family']

Current parameters without a versioned optimization rule

  • abort_on_priority_when_disabled
  • admin_api_key
  • allow_auto_truncate
  • allowed_media_domains
  • api_key
  • asr_max_buffer_seconds
  • asr_max_concurrent_sessions
  • attn_cp_size
  • base_gpu_id
  • batch_notify_size
  • bf16_gemm_backend
  • bucket_e2e_request_latency
  • bucket_inter_token_latency
  • bucket_time_to_first_token
  • c128_page_size
  • chat_template
  • checkpoint_engine_wait_weights_before_ready
  • completion_template
  • config
  • constrained_json_disable_any_whitespace
  • constrained_json_whitespace_pattern
  • context_length
  • cp_strategy
  • cpu_offload_gb
  • crash_dump_folder
  • cuda_graph_backend_decode
  • cuda_graph_config
  • custom_sigquit_handler
  • custom_weight_loader
  • dcp_comm_backend
  • dcp_replicate_q_proj
  • dcp_size
  • debug_cuda_graph
  • debug_tensor_dump_input_file
  • debug_tensor_dump_layers
  • debug_tensor_dump_output_folder
  • decode_log_interval
  • decoupled_spec_bind_endpoint
  • decoupled_spec_connect_endpoints
  • decoupled_spec_rank
  • decoupled_spec_role
  • decrypted_config_file
  • decrypted_draft_config_file
  • deepep_config
  • deepep_dispatcher_output_dtype
  • deepep_mode
  • deepep_v2_mode
  • default_chat_template_kwargs
  • default_priority_value
  • delete_ckpt_after_loading
  • detokenizer_worker_num
  • device
  • disable_attn_tp_gather
  • disable_chunked_prefix_cache
  • disable_cuda_graph_padding
  • disable_decode_cuda_graph
  • disable_flashinfer_autotune
  • disable_flashinfer_cutlass_moe_fp4_allgather
  • disable_hybrid_swa_memory
  • disable_outlines_disk_cache
  • disable_prefill_cuda_graph
  • disable_priority_preemption
  • disable_shared_experts_fusion
  • disable_tokenizer_batch_decode
  • disaggregation_bootstrap_port
  • disaggregation_decode_enable_offload_kvcache
  • disaggregation_decode_enable_radix_cache
  • disaggregation_decode_extra_slots
  • disaggregation_decode_polling_interval
  • disaggregation_decode_retraction_backup
  • disaggregation_ib_device
  • disaggregation_mode
  • disaggregation_transfer_backend
  • dist_init_addr
  • dist_timeout
  • dllm_algorithm
  • dllm_algorithm_config
  • dllm_fdfo
  • download_dir
  • dp_size
  • dsa_paged_mqa_logits_backend
  • dsa_topk_backend
  • dsv4_prefill_backend
  • dtype
  • dwdp_size
  • dynamic_batch_tokenizer_batch_size
  • dynamic_batch_tokenizer_batch_timeout
  • elastic_ep_backend
  • elastic_ep_initial_size
  • elastic_ep_scale_timeout
  • enable_adaptive_dispatch_to_encoder
  • enable_aiter_allreduce_fusion
  • enable_attn_tp_input_scattered
  • enable_broadcast_mm_inputs_process
  • enable_cache_report
  • enable_cp_decode_attn_tp
  • enable_cudagraph_gc
  • enable_custom_logit_processor
  • enable_deepseek_v4_fp4_indexer
  • enable_dense_mlp_attn_tp
  • enable_deterministic_inference
  • enable_dp_attention_local_control_broadcast
  • enable_dp_lm_head
  • enable_draft_weights_cpu_backup
  • enable_dsa_cache_layer_split
  • enable_dynamic_batch_tokenizer
  • enable_dynamic_chunking
  • enable_elastic_expert_backup
  • enable_eplb
  • enable_expert_distribution_metrics
  • enable_flexkv
  • enable_forward_pass_metrics
  • enable_fp32_lm_head
  • enable_fused_moe_sum_all_reduce
  • enable_fused_qk_norm_rope
  • enable_hierarchical_cache
  • enable_hisparse
  • enable_http2
  • enable_int8_mamba_checkpoint
  • enable_layerwise_nvtx_marker
  • enable_lean_attention
  • enable_linear_replayssm
  • enable_lmcache
  • enable_lora
  • enable_lora_overlap_loading
  • enable_mamba_cache_stochastic_rounding
  • enable_memory_saver
  • enable_metrics
  • enable_metrics_for_all_schedulers
  • enable_mfu_metrics
  • enable_mis
  • enable_mm_global_cache
  • enable_multi_layer_eagle
  • enable_multimodal
  • enable_nccl_nvls
  • enable_p2p_check
  • enable_page_major_kv_layout
  • enable_pdmux
  • enable_precise_embedding_interpolation
  • enable_prefill_cp
  • enable_prefill_delayer
  • enable_prefix_mm_cache
  • enable_priority_scheduling
  • enable_profile_cuda_graph
  • enable_quant_communications
  • enable_request_time_stats_logging
  • enable_return_hidden_states
  • enable_return_indexer_topk
  • enable_return_routed_experts
  • enable_scattered_sconv
  • enable_session_radix_cache
  • enable_shared_experts_attn_tp
  • enable_single_batch_overlap
  • enable_ssl_refresh
  • enable_streaming_session
  • enable_strict_thinking
  • enable_symm_mem
  • enable_tf32_matmul
  • enable_tokenizer_batch_encode
  • enable_torch_compile
  • enable_torch_compile_debug_mode
  • enable_torch_symm_mem
  • enable_tp_lm_head_all_to_all
  • enable_trace
  • enable_two_batch_overlap
  • enable_unified_memory
  • enable_w4a4_mxfp4_megamoe
  • enable_waterfill
  • enable_weights_cpu_backup
  • encoder_bootstrap_port
  • encoder_only
  • encoder_register_urls
  • encoder_transfer_backend
  • encoder_urls
  • enforce_disable_flashinfer_allreduce_fusion
  • enforce_shared_experts_fusion
  • engine_info_bootstrap_port
  • ep_dispatch_algorithm
  • ep_join_mode
  • ep_join_rank_offset
  • ep_num_redundant_experts
  • eplb_algorithm
  • eplb_min_rebalancing_utilization_threshold
  • eplb_rebalance_layers_per_chunk
  • eplb_rebalance_num_iterations
  • expert_balancedness_report_mode
  • expert_distribution_recorder_buffer_size
  • expert_distribution_recorder_mode
  • experts_shared_outer_loras
  • export_metrics_to_file
  • export_metrics_to_file_dir
  • extra_metric_labels
  • fastapi_root_path
  • file_storage_path
  • flashinfer_allreduce_fusion_backend
  • flashinfer_autotune_skip_ops
  • flashinfer_mla_disable_ragged
  • flashinfer_mxfp4_moe_precision
  • flexkv_config_file
  • forward_hooks
  • forward_pass_metrics_ipc_name
  • forward_pass_metrics_worker_id
  • fp4_gemm_runner_backend
  • fp8_gemm_runner_backend
  • fuseep_mode
  • gated_launch_port
  • gc_threshold
  • gc_warning_threshold_secs
  • generation_tokens_buckets
  • gpu_id_step
  • grammar_backend
  • grpc_port
  • hf_chat_template_name
  • hicache_host_memory_mode
  • hicache_io_backend
  • hicache_mem_layout
  • hicache_ratio
  • hicache_size
  • hicache_storage_backend
  • hicache_storage_backend_extra_config
  • hicache_storage_prefetch_policy
  • hicache_storage_prefetch_retry_max_attempts
  • hicache_storage_prefetch_retry_poll_interval
  • hicache_write_policy
  • hisparse_config
  • host
  • http2_initial_connection_window_size
  • http2_max_concurrent_streams
  • image_processor_backend
  • init_expert_location
  • int8_mamba_ckpt_size
  • is_embedding
  • json_model_override_args
  • kt_cpuinfer
  • kt_max_deferred_experts_per_token
  • kt_method
  • kt_num_gpu_experts
  • kt_threadpool_count
  • kt_weight_path
  • kv_cache_dtype
  • kv_canary
  • kv_canary_real_data
  • kv_canary_sweep_interval
  • kv_events_config
  • language_model_only
  • language_only
  • limit_mm_data_per_request
  • linear_attn_backend
  • linear_attn_verify_backend
  • linear_replayssm_cache_len
  • lmcache_config_file
  • load_balance_method
  • load_format
  • load_publish_endpoint
  • load_snapshot_publish_interval
  • log_level
  • log_level_http
  • log_requests
  • log_requests_format
  • log_requests_level
  • log_requests_target
  • lora_backend
  • lora_drain_wait_threshold
  • lora_eviction_policy
  • lora_paths
  • lora_strict_loading
  • lora_target_modules
  • lora_use_virtual_experts
  • mamba_backend
  • mamba_cache_philox_rounds
  • mamba_max_states_per_path
  • mamba_track_interval
  • max_ep_size
  • max_loaded_loras
  • max_lora_chunk_size
  • max_lora_rank
  • max_loras_per_batch
  • max_mamba_cache_size
  • max_queued_requests
  • max_running_requests
  • max_total_tokens
  • media_url_max_file_size_mb
  • min_free_slots_delay
  • mlx_enable_sampling
  • mm_attention_backend
  • mm_enable_dp_encoder
  • mm_feature_transport
  • mm_global_cache_backend
  • mm_io_worker_num
  • mm_preprocess_cache_size_mb
  • mm_process_config
  • mm_processor_worker_num
  • model_checksum
  • model_config_parser
  • model_impl
  • model_loader_extra_config
  • model_path
  • modelexpress_config
  • modelopt_checkpoint_restore_path
  • modelopt_checkpoint_save_path
  • modelopt_export_path
  • modelopt_quant
  • moe_a2a_backend
  • moe_dense_tp_size
  • moe_dp_size
  • mooncake_ib_device
  • msprobe_dump_config
  • nccl_port
  • nnodes
  • node_rank
  • num_reserved_decode_tokens
  • numa_node
  • offload_group_size
  • offload_mode
  • offload_num_in_group
  • offload_prefetch_step
  • optimistic_prefill_attempts
  • otlp_traces_endpoint
  • pdmux_config_path
  • port
  • pp_async_batch_depth
  • pp_max_micro_batch_size
  • pp_size
  • pre_warm_nccl
  • preferred_sampling_params
  • prefill_decode_interval
  • prefill_delayer_forward_passes_buckets
  • prefill_delayer_max_delay_ms
  • prefill_delayer_max_delay_passes
  • prefill_delayer_queue_min_ratio
  • prefill_delayer_token_usage_low_watermark
  • prefill_delayer_wait_seconds_buckets
  • prefill_max_requests
  • prefill_only_disable_kv_cache
  • priority_scheduling_preemption_threshold
  • prompt_tokens_buckets
  • quantization
  • quantization_param_path
  • quantize_and_serve
  • radix_cache_backend
  • radix_eviction_policy
  • random_seed
  • reasoning_parser
  • remote_instance_weight_loader_backend
  • remote_instance_weight_loader_seed_instance_ip
  • remote_instance_weight_loader_seed_instance_service_port
  • remote_instance_weight_loader_send_weights_group_ports
  • remote_instance_weight_loader_start_seed_via_transfer_engine
  • retraction_policy
  • return_hidden_states_mode
  • revision
  • rl_on_policy_target
  • rl_quant_profile
  • sampling_backend
  • sampling_defaults
  • schedule_low_priority_values_first
  • served_model_name
  • show_time_cost
  • sidecar
  • sidecar_args
  • skip_server_warmup
  • skip_tokenizer_init
  • sleep_on_idle
  • sm_group_num
  • smg_http_sidecar_port
  • soft_watchdog_timeout
  • spec_trace_dir
  • speculative_accept_threshold_acc
  • speculative_accept_threshold_single
  • speculative_adaptive
  • speculative_adaptive_config
  • speculative_attention_mode
  • speculative_dflash_block_size
  • speculative_draft_attention_backend
  • speculative_draft_kv_cache_dtype
  • speculative_draft_load_format
  • speculative_draft_model_path
  • speculative_draft_model_quantization
  • speculative_draft_model_revision
  • speculative_draft_window_size
  • speculative_dsa_topk_backend
  • speculative_dspark_align_verify_tokens_to_graph_tier
  • speculative_dspark_block_size
  • speculative_dspark_confidence_sts_path
  • speculative_dspark_sps_table_path
  • speculative_eagle_topk
  • speculative_moe_a2a_backend
  • speculative_moe_runner_backend
  • speculative_ngram_capacity
  • speculative_ngram_external_corpus_max_tokens
  • speculative_ngram_external_corpus_path
  • speculative_ngram_external_sam_budget
  • speculative_ngram_match_type
  • speculative_ngram_max_bfs_breadth
  • speculative_ngram_max_trie_depth
  • speculative_ngram_min_bfs_breadth
  • speculative_skip_dp_mlp_sync
  • speculative_token_map
  • speculative_use_rejection_sampling
  • ssl_ca_certs
  • ssl_certfile
  • ssl_keyfile
  • ssl_keyfile_password
  • startup_weight_load_mode
  • stat_loggers
  • stream_interval
  • stream_response_default_include_usage
  • strip_thinking_cache
  • swa_full_tokens_ratio
  • tbo_token_distribution_threshold
  • tokenizer_backend
  • tokenizer_metrics_allowed_custom_labels
  • tokenizer_metrics_custom_labels_header
  • tokenizer_mode
  • tokenizer_path
  • tokenizer_worker_num
  • tool_call_parser
  • tool_server
  • torch_compile_max_bs
  • trace_modules
  • triton_attention_num_kv_splits
  • triton_attention_reduce_in_fp32
  • triton_attention_split_tile_size
  • trust_mm_content_hashes
  • trust_remote_code
  • use_ray
  • uvicorn_access_log_exclude_prefixes
  • warmups
  • watchdog_timeout
  • weight_cache_mode
  • weight_cache_socket
  • weight_cache_timeout
  • weight_loader_disable_mmap
  • weight_loader_drop_cache_after_load
  • weight_loader_prefetch_checkpoints
  • weight_loader_prefetch_num_threads
  • weight_version

Unknown parameters are not automatically trusted. Run InferOpt semantic/safety analysis and real benchmarks before promoting them to validated rules.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    sglang-compatibilityTracks compatibility with evolving SGLang ServerArgs

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions