版本:开源版5.8.9, 部署方式:3节点core集群。
日志:
2026-09-15T01:28:51.696099+08:00 [info] clientid: universal_ies832.NBDL_DQ_A1_Gateway_002|securemode=1,authType=1,signmethod=HmacSHA256,timestamp=1755150868383,tenan
tId=1919995136244723714|, msg: terminate, peername: 192.168.2.130:44930, username: NBDL_DQ_A1_Gateway_002&universal_ies832, reason: {shutdown,keepalive_timeout}
2026-09-15T02:44:40.720951+08:00 [info] clientid: universal_ies832.BLTY_GateWay_001|securemode=1,authType=1,signmethod=HmacSHA256,timestamp=1757298449161,tenantId=19
19995136244723714|, msg: terminate, peername: 192.168.2.130:41884, username: BLTY_GateWay_001&universal_ies832, reason: {shutdown,keepalive_timeout}
2026-09-15T03:07:22.320107+08:00 [info] clientid: universal_ies832.NBDL_DQ_A1_Gateway_002|securemode=1,authType=1,signmethod=HmacSHA256,timestamp=1755150868383,tenan
tId=1919995136244723714|, msg: terminate, peername: 192.168.2.130:46186, username: NBDL_DQ_A1_Gateway_002&universal_ies832, reason: {shutdown,keepalive_timeout}
2026-09-15T03:07:46.068636+08:00 [info] clientid: universal_ies832.NBDL_DQ_A1_Gateway_002|securemode=1,authType=1,signmethod=HmacSHA256,timestamp=1755150868383,tenan
tId=1919995136244723714|, msg: terminate, peername: 192.168.2.130:63666, username: NBDL_DQ_A1_Gateway_002&universal_ies832, reason: {shutdown,tcp_closed}
2026-09-15T04:07:46.730471+08:00 [info] clientid: universal_ies832.NBDL_DQ_A1_Gateway_002|securemode=1,authType=1,signmethod=HmacSHA256,timestamp=1755150868383,tenan
tId=1919995136244723714|, msg: terminate, peername: 192.168.2.130:63676, username: NBDL_DQ_A1_Gateway_002&universal_ies832, reason: {shutdown,tcp_closed}
2026-09-15T07:35:11.981518+08:00 [info] Mria(Membership): Node emqx@192.168.2.131 down
2026-09-15T07:35:11.982054+08:00 [warning] msg: cm_registry_node_down, node: ‘emqx@192.168.2.131’
2026-09-15T07:35:12.046948+08:00 [info] Mria(Membership): Mnesia emqx@192.168.2.131 down
2026-09-15T07:35:12.059106+08:00 [warning] msg: cm_registry_mnesia_down, node: ‘emqx@192.168.2.131’
2026-09-15T07:35:12.708508+08:00 [info] Mria(Membership): Node emqx@192.168.2.131 up
2026-09-15T07:35:12.709836+08:00 [warning] msg: exclusive_sub_worker_unexpected_info, info: {mnesia_locker,‘emqx@192.168.2.131’,granted}
2026-09-15T07:35:12.710022+08:00 [error] msg: unexpected_info, info: {mnesia_locker,‘emqx@192.168.2.131’,granted}
2026-09-15T07:35:12.710677+08:00 [critical] msg: Core cluster partition, context: running_partitioned_network, from: ‘emqx@192.168.2.131’
2026-09-15T07:35:12.710667+08:00 [error] Mnesia(‘emqx@192.168.2.130’): ** ERROR ** mnesia_event got {inconsistent_database, running_partitioned_network, ‘emqx@192.16
8.2.131’}
2026-09-15T07:35:12.711090+08:00 [warning] msg: alarm_is_activated, message: <<“Partition occurs at node emqx@192.168.2.131”>>, name: partition
2026-09-15T07:35:14.985347+08:00 [info] msg: mria_node_monitor_confirm, status: up, target_node: ‘emqx@192.168.2.131’
2026-09-15T07:35:30.718879+08:00 [notice] msg: Mria is restarting to join the cluster, seed: ‘emqx@192.168.2.132’
2026-09-15T07:35:30.719792+08:00 [warning] msg: Stopping mria, reason: heal
2026-09-15T07:35:30.720065+08:00 [notice] msg: stopping_emqx_apps
2026-09-15T07:35:30.766199+08:00 [notice] Application: quicer. Exited: stopped. Type: temporary.
2026-09-15T07:35:30.772166+08:00 [notice] Application: emqx_connector_jwt. Exited: stopped. Type: temporary.
2026-09-15T07:35:30.779391+08:00 [notice] Application: emqx_psk. Exited: stopped. Type: temporary.
2026-09-15T07:35:30.785414+08:00 [notice] Application: emqx_telemetry. Exited: stopped. Type: temporary.
2026-09-15T07:35:30.785945+08:00 [notice] Application: emqx_gateway_mqttsn. Exited: stopped. Type: temporary.
2026-09-15T07:35:30.786283+08:00 [notice] Application: emqx_gateway_lwm2m. Exited: stopped. Type: temporary.
2026-09-15T07:35:30.786457+08:00 [notice] Application: emqx_gateway_coap. Exited: stopped. Type: temporary.
2026-09-15T07:35:30.786614+08:00 [notice] Application: emqx_gateway_stomp. Exited: stopped. Type: temporary.
2026-09-15T07:35:30.786978+08:00 [notice] Application: emqx_gateway_exproto. Exited: stopped. Type: temporary.
2026-09-15T07:35:30.794057+08:00 [notice] Application: emqx_gateway. Exited: stopped. Type: temporary.
2026-09-15T07:35:30.799614+08:00 [notice] Application: emqx_auth_mysql. Exited: stopped. Type: temporary.
2026-09-15T07:35:30.799976+08:00 [notice] Application: emqx_mysql. Exited: stopped. Type: temporary.
2026-09-15T07:35:30.805934+08:00 [notice] Application: emqx_prometheus. Exited: stopped. Type: temporary.
2026-09-15T07:35:30.811385+08:00 [notice] Application: emqx_auth_mnesia. Exited: stopped. Type: temporary.
2026-09-15T07:35:30.817364+08:00 [notice] Application: emqx_auth_redis. Exited: stopped. Type: temporary.
2026-09-15T07:35:30.823712+08:00 [notice] Application: emqx_auth_http. Exited: stopped. Type: temporary.
2026-09-15T07:35:30.830293+08:00 [notice] Application: emqx_auth_mongodb. Exited: stopped. Type: temporary.
2026-09-15T07:35:30.835981+08:00 [notice] Application: emqx_auth_jwt. Exited: stopped. Type: temporary.
2026-09-15T07:35:30.841852+08:00 [notice] Application: emqx_auth_ldap. Exited: stopped. Type: temporary.
2026-09-15T07:35:30.847862+08:00 [notice] Application: emqx_auth_postgresql. Exited: stopped. Type: temporary.
2026-09-15T07:35:30.852310+08:00 [info] msg: stopping_http_connector, connector: <<“emqx_authn_http:64516”>>
2026-09-15T07:35:30.871149+08:00 [notice] Application: emqx_auth. Exited: stopped. Type: temporary.
2026-09-15T07:35:30.877198+08:00 [notice] Application: emqx_auto_subscribe. Exited: stopped. Type: temporary.
2026-09-15T07:35:30.886367+08:00 [notice] Application: emqx_dashboard. Exited: stopped. Type: temporary.
2026-09-15T07:35:30.904529+08:00 [notice] Application: emqx_rule_engine. Exited: stopped. Type: temporary.
2026-09-15T07:35:30.919301+08:00 [notice] Application: emqx_bridge. Exited: stopped. Type: temporary.
2026-09-15T07:35:30.920397+08:00 [notice] Application: emqx_ldap. Exited: stopped. Type: temporary.
2026-09-15T07:35:30.927153+08:00 [notice] Application: emqx_connector. Exited: stopped. Type: temporary.
2026-09-15T07:35:30.927721+08:00 [notice] Application: emqx_mongodb. Exited: stopped. Type: temporary.
2026-09-15T07:35:30.928356+08:00 [info] msg: exhook_mgr_terminated, reason: shutdown, servers: #{}
2026-09-15T07:35:30.934275+08:00 [notice] Application: emqx_exhook. Exited: stopped. Type: temporary.
2026-09-15T07:35:30.957399+08:00 [notice] Application: emqx_retainer. Exited: stopped. Type: temporary.
2026-09-15T07:35:30.976163+08:00 [notice] Application: emqx_modules. Exited: stopped. Type: temporary.
2026-09-15T07:35:30.982446+08:00 [notice] Application: emqx_slow_subs. Exited: stopped. Type: temporary.
2026-09-15T07:35:30.988411+08:00 [notice] Application: emqx_management. Exited: stopped. Type: temporary.
2026-09-15T07:35:30.994389+08:00 [notice] Application: emqx_plugins. Exited: stopped. Type: temporary.
2026-09-15T07:35:30.994928+08:00 [notice] Application: emqx_bridge_mqtt. Exited: stopped. Type: temporary.
2026-09-15T07:35:30.996541+08:00 [notice] Application: emqx_bridge_http. Exited: stopped. Type: temporary.
2026-09-15T07:35:30.996830+08:00 [notice] Application: emqx_redis. Exited: stopped. Type: temporary.
2026-09-15T07:35:30.997059+08:00 [notice] Application: emqx_postgresql. Exited: stopped. Type: temporary.
2026-09-15T07:35:31.014404+08:00 [notice] Application: emqx_resource. Exited: stopped. Type: temporary.
2026-09-15T07:35:31.016106+08:00 [notice] tcp:default stopped on 0.0.0.0:1883
2026-09-15T07:35:31.020814+08:00 [notice] ssl:default stopped on 0.0.0.0:8883
2026-09-15T07:35:31.100237+08:00 [notice] Application: emqx. Exited: stopped. Type: permanent.
2026-09-15T07:35:31.101037+08:00 [notice] Application: emqx_ds_backends. Exited: stopped. Type: temporary.
2026-09-15T07:35:31.115741+08:00 [notice] Application: emqx_durable_storage. Exited: stopped. Type: temporary.
2026-09-15T07:35:31.116161+08:00 [notice] Application: emqx_utils. Exited: stopped. Type: temporary.
2026-09-15T07:35:31.126884+08:00 [notice] Application: cowboy. Exited: stopped. Type: temporary.
2026-09-15T07:35:31.140147+08:00 [notice] Application: ranch. Exited: stopped. Type: temporary.
2026-09-15T07:35:31.155944+08:00 [notice] Application: bcrypt. Exited: stopped. Type: temporary.
2026-09-15T07:35:31.156355+08:00 [notice] Application: emqx_http_lib. Exited: stopped. Type: temporary.
2026-09-15T07:35:31.166931+08:00 [notice] Application: gproc. Exited: stopped. Type: temporary.
2026-09-15T07:35:31.181337+08:00 [notice] Application: esockd. Exited: stopped. Type: permanent.
2026-09-15T07:35:31.196590+08:00 [notice] Application: emqx_conf. Exited: stopped. Type: temporary.
2026-09-15T07:35:31.209492+08:00 [notice] Application: ekka. Exited: stopped. Type: temporary.
2026-09-15T07:35:31.253510+08:00 [notice] msg: Mria is stopped
2026-09-15T07:35:31.264093+08:00 [notice] Application: mria. Exited: stopped. Type: temporary.
2026-09-15T07:35:31.295314+08:00 [notice] Application: mnesia. Exited: stopped. Type: temporary.
2026-09-15T07:35:32.418414+08:00 [notice] msg: Starting mria
2026-09-15T07:35:32.420950+08:00 [notice] msg: Starting mnesia
2026-09-15T07:35:32.426000+08:00 [notice] msg: Starting shards
2026-09-15T07:35:32.750320+08:00 [info] msg: Setting RLOG shard config, tables: [‘$mria_rlog_sync’,mria_schema], shard: ‘$mria_meta_shard’
2026-09-15T07:35:32.751141+08:00 [info] msg: Converging schema
2026-09-15T07:35:32.753221+08:00 [info] msg: Setting RLOG shard config, tables: [‘$mria_rlog_sync’,mria_schema], shard: ‘$mria_meta_shard’
2026-09-15T07:35:32.755023+08:00 [info] msg: Setting RLOG shard config, tables: [bpapi,emqx_app,emqx_banned,emqx_banned_rules,emqx_trace], shard: emqx_common_shard
2026-09-15T07:35:32.756781+08:00 [info] msg: Setting RLOG shard config, tables: [cluster_rpc_commit,cluster_rpc_mfa], shard: emqx_cluster_rpc_shard
2026-09-15T07:35:32.758337+08:00 [info] msg: Setting RLOG shard config, tables: [cluster_rpc_commit,cluster_rpc_mfa], shard: emqx_cluster_rpc_shard
2026-09-15T07:35:32.759630+08:00 [info] msg: Setting RLOG shard config, tables: [emqx_acl], shard: emqx_acl_sharded
2026-09-15T07:35:32.760858+08:00 [info] msg: Setting RLOG shard config, tables: [emqx_activated_alarm,emqx_dashboard_monitor,emqx_deactivated_alarm,emqx_delayed,emqx
_ds_builtin_local_metadata_tab,emqx_ds_builtin_local_timestamp_tab], shard: undefined
2026-09-15T07:35:32.762271+08:00 [info] msg: Setting RLOG shard config, tables: [emqx_admin,emqx_admin_jwt], shard: emqx_dashboard_shard
2026-09-15T07:35:32.763673+08:00 [info] msg: Setting RLOG shard config, tables: [emqx_admin,emqx_admin_jwt], shard: emqx_dashboard_shard
2026-09-15T07:35:32.765308+08:00 [info] msg: Setting RLOG shard config, tables: [bpapi,emqx_app,emqx_banned,emqx_banned_rules,emqx_trace], shard: emqx_common_shard
2026-09-15T07:35:32.766687+08:00 [info] msg: Setting RLOG shard config, tables: [emqx_authn_mnesia,emqx_authn_scram_mnesia], shard: emqx_authn_shard
2026-09-15T07:35:32.768086+08:00 [info] msg: Setting RLOG shard config, tables: [emqx_authn_mnesia,emqx_authn_scram_mnesia], shard: emqx_authn_shard
2026-09-15T07:35:32.769597+08:00 [info] msg: Setting RLOG shard config, tables: [bpapi,emqx_app,emqx_banned,emqx_banned_rules,emqx_trace], shard: emqx_common_shard
2026-09-15T07:35:32.771464+08:00 [info] msg: Setting RLOG shard config, tables: [bpapi,emqx_app,emqx_banned,emqx_banned_rules,emqx_trace], shard: emqx_common_shard
2026-09-15T07:35:32.772918+08:00 [info] msg: Setting RLOG shard config, tables: [emqx_channel_registry], shard: emqx_cm_shard
2026-09-15T07:35:32.774947+08:00 [info] msg: Setting RLOG shard config, tables: [emqx_activated_alarm,emqx_dashboard_monitor,emqx_deactivated_alarm,emqx_delayed,emqx
_ds_builtin_local_metadata_tab,emqx_ds_builtin_local_timestamp_tab], shard: undefined
2026-09-15T07:35:32.776438+08:00 [info] msg: Setting RLOG shard config, tables: [emqx_activated_alarm,emqx_dashboard_monitor,emqx_deactivated_alarm,emqx_delayed,emqx
_ds_builtin_local_metadata_tab,emqx_ds_builtin_local_timestamp_tab], shard: undefined
2026-09-15T07:35:32.777863+08:00 [info] msg: Setting RLOG shard config, tables: [emqx_activated_alarm,emqx_dashboard_monitor,emqx_deactivated_alarm,emqx_delayed,emqx
_ds_builtin_local_metadata_tab,emqx_ds_builtin_local_timestamp_tab], shard: undefined
2026-09-15T07:35:32.780480+08:00 [info] msg: Setting RLOG shard config, tables: [emqx_activated_alarm,emqx_dashboard_monitor,emqx_deactivated_alarm,emqx_delayed,emqx
_ds_builtin_local_metadata_tab,emqx_ds_builtin_local_timestamp_tab], shard: undefined
2026-09-15T07:35:32.782118+08:00 [info] msg: Setting RLOG shard config, tables: [emqx_activated_alarm,emqx_dashboard_monitor,emqx_deactivated_alarm,emqx_delayed,emqx
_ds_builtin_local_metadata_tab,emqx_ds_builtin_local_timestamp_tab], shard: undefined
2026-09-15T07:35:32.783405+08:00 [info] msg: Setting RLOG shard config, tables: [emqx_exclusive_subscription,emqx_exclusive_subscription_v2], shard: emqx_exclusive_s
hard
2026-09-15T07:35:32.784716+08:00 [info] msg: Setting RLOG shard config, tables: [emqx_exclusive_subscription,emqx_exclusive_subscription_v2], shard: emqx_exclusive_s
hard
2026-09-15T07:35:32.786511+08:00 [info] msg: Setting RLOG shard config, tables: [emqx_psk], shard: emqx_psk_shard
2026-09-15T07:35:32.788761+08:00 [info] msg: Setting RLOG shard config, tables: [emqx_retainer_index,emqx_retainer_index_meta,emqx_retainer_message], shard: emqx_ret
ainer_shard
2026-09-15T07:35:32.790795+08:00 [info] msg: Setting RLOG shard config, tables: [emqx_retainer_index,emqx_retainer_index_meta,emqx_retainer_message], shard: emqx_ret
ainer_shard
2026-09-15T07:35:32.793113+08:00 [info] msg: Setting RLOG shard config, tables: [emqx_retainer_index,emqx_retainer_index_meta,emqx_retainer_message], shard: emqx_ret
ainer_shard
2026-09-15T07:35:32.794964+08:00 [info] msg: Setting RLOG shard config, tables: [emqx_route,emqx_route_filters,emqx_routing_node,emqx_trie], shard: route_shard
2026-09-15T07:35:32.799455+08:00 [info] msg: Setting RLOG shard config, tables: [emqx_route,emqx_route_filters,emqx_routing_node,emqx_trie], shard: route_shard
2026-09-15T07:35:32.804643+08:00 [info] msg: Setting RLOG shard config, tables: [emqx_route,emqx_route_filters,emqx_routing_node,emqx_trie], shard: route_shard
2026-09-15T07:35:32.806478+08:00 [info] msg: Setting RLOG shard config, tables: [emqx_shared_subscription], shard: emqx_shared_sub_shard
2026-09-15T07:35:32.808100+08:00 [info] msg: Setting RLOG shard config, tables: [emqx_telemetry], shard: emqx_telemetry_shard
2026-09-15T07:35:32.810504+08:00 [info] msg: Setting RLOG shard config, tables: [bpapi,emqx_app,emqx_banned,emqx_banned_rules,emqx_trace], shard: emqx_common_shard
2026-09-15T07:35:32.812222+08:00 [info] msg: Setting RLOG shard config, tables: [emqx_route,emqx_route_filters,emqx_routing_node,emqx_trie], shard: route_shard
2026-09-15T07:35:32.813465+08:00 [info] Mria(Membership): Node emqx@192.168.2.132 up
2026-09-15T07:35:32.813902+08:00 [info] msg: starting_rlog_shard, shard: ‘$mria_meta_shard’
2026-09-15T07:35:32.814677+08:00 [info] Mria(Membership): Node emqx@192.168.2.131 up
2026-09-15T07:35:32.815562+08:00 [notice] msg: Mria is running
2026-09-15T07:35:32.815590+08:00 [info] msg: Shard fully up, node: ‘emqx@192.168.2.130’, shard: ‘$mria_meta_shard’
2026-09-15T07:35:32.816013+08:00 [info] msg: Starting ekka
2026-09-15T07:35:32.816932+08:00 [notice] msg: Mria has joined the cluster, status: #{members => [{member,‘emqx@192.168.2.130’,undefined,<<0,6,91,121,227,228,32,37,1
56,25,0,11,150,197,0,0>>,3325906226,up,running,{1789,428932,812889},core},{member,‘emqx@192.168.2.132’,undefined,undefined,undefined,up,running,{1789,428932,814655},
core}],running_nodes => [‘emqx@192.168.2.130’,‘emqx@192.168.2.131’,‘emqx@192.168.2.132’],rlog => #{role => core,backend => rlog,imbalance => },partitions => ,reb
alance_status => not_started,stopped_nodes => }, seed: ‘emqx@192.168.2.132’
2026-09-15T07:35:32.817923+08:00 [info] msg: Ekka is running
2026-09-15T07:35:32.818251+08:00 [notice] msg: (re)starting_emqx_apps
2026-09-15T07:35:32.830257+08:00 [info] msg: starting_rlog_shard, shard: emqx_cluster_rpc_shard
2026-09-15T07:35:32.831438+08:00 [info] msg: wait_for_cluster_rpc_shard, result: ok
2026-09-15T07:35:32.831530+08:00 [info] msg: Shard fully up, node: ‘emqx@192.168.2.130’, shard: emqx_cluster_rpc_shard
2026-09-15T07:35:32.831783+08:00 [info] msg: wait_for_cluster_rpc_tables, result: ok
2026-09-15T07:35:32.844566+08:00 [info] msg: sync_cluster_conf_success, data_dir: data, has_deprecated_file: false, tnx_id: 12, local_release: v5.8.9, remote_release
: v5.8.9, synced_from_node: ‘emqx@192.168.2.131’
2026-09-15T07:35:33.189582+08:00 [info] msg: starting_rlog_shard, shard: emqx_common_shard
2026-09-15T07:35:33.190482+08:00 [info] msg: Shard fully up, node: ‘emqx@192.168.2.130’, shard: emqx_common_shard
2026-09-15T07:35:33.245725+08:00 [info] msg: starting_rlog_shard, shard: route_shard
2026-09-15T07:35:33.247314+08:00 [info] msg: Shard fully up, node: ‘emqx@192.168.2.130’, shard: route_shard
2026-09-15T07:35:33.248274+08:00 [info] msg: routing_schema_used, schema: v2
2026-09-15T07:35:33.265367+08:00 [info] msg: starting_rlog_shard, shard: emqx_shared_sub_shard
2026-09-15T07:35:33.266478+08:00 [info] msg: Shard fully up, node: ‘emqx@192.168.2.130’, shard: emqx_shared_sub_shard
2026-09-15T07:35:33.304639+08:00 [info] msg: starting_rlog_shard, shard: emqx_exclusive_shard
2026-09-15T07:35:33.305567+08:00 [info] msg: Shard fully up, node: ‘emqx@192.168.2.130’, shard: emqx_exclusive_shard
2026-09-15T07:35:33.317769+08:00 [info] msg: starting_rlog_shard, shard: emqx_cm_shard
2026-09-15T07:35:33.318900+08:00 [info] msg: Shard fully up, node: ‘emqx@192.168.2.130’, shard: emqx_cm_shard
2026-09-15T07:35:33.362515+08:00 [info] msg: wait_for_cluster_rpc_shard, result: ok
2026-09-15T07:35:33.363060+08:00 [info] msg: wait_for_cluster_rpc_tables, result: ok
2026-09-15T07:35:33.364225+08:00 [info] msg: CMD_overridden, cmd: observer, mf: {emqx_observer_cli,cmd}
2026-09-15T07:35:33.364660+08:00 [info] msg: CMD_overridden, cmd: cluster_call, mf: {emqx_conf_cli,admins}
2026-09-15T07:35:33.365011+08:00 [info] msg: CMD_overridden, cmd: conf, mf: {emqx_conf_cli,conf}
2026-09-15T07:35:33.373168+08:00 [info] msg: starting_rlog_shard, shard: emqx_retainer_shard
2026-09-15T07:35:33.374862+08:00 [info] msg: Shard fully up, node: ‘emqx@192.168.2.130’, shard: emqx_retainer_shard
2026-09-15T07:35:33.403933+08:00 [info] msg: starting_rlog_shard, shard: emqx_dashboard_shard
2026-09-15T07:35:33.405299+08:00 [info] msg: Shard fully up, node: ‘emqx@192.168.2.130’, shard: emqx_dashboard_shard
2026-09-15T07:35:33.408466+08:00 [info] msg: loading_desc, file: /opt/emqx/lib/emqx_dashboard-5.2.3/priv/desc.en.hocon
2026-09-15T07:35:33.632521+08:00 [info] msg: loading_desc, file: /opt/emqx/lib/emqx_dashboard-5.2.3/priv/desc.zh.hocon
2026-09-15T07:35:33.849544+08:00 [info] msg: started_listener_ok, name: ‘http:dashboard’, pid: <0.760398.0>, port: 18083
2026-09-15T07:35:33.992616+08:00 [info] msg: regenerate_dispatch, listeners: [‘http:dashboard’], i18n_lang: en, elapsed_ms: 142
2026-09-15T07:35:34.008294+08:00 [info] msg: starting_http_connector, config: #{ssl => #{depth => 10,verify => verify_peer,hibernate_after => 5000,enable => false,ci
phers => ,log_level => notice,versions => [‘tlsv1.3’,‘tlsv1.2’],secure_renegotiate => true,reuse_sessions => true},connect_timeout => 15000,mechanism => password_b
ased,pool_size => 8,enable => true,body => #{password => <<“${password}”>>,username => <<“${username}”>>,clientid => <<“${clientid}”>>,peerhost => <<“${peerhost}”>>}
,headers => #{<<“content-type”>> => <<“application/json”>>},url => <<“http://192.168.2.130:9001/v1/mqtt/auth”>>,method => post,backend => http,request_timeout => 500
0,request_base => #{port => 9001,scheme => http,host => {192,168,2,130}},pool_type => random,enable_pipelining => 100}, connector: <<“emqx_authn_http:125427”>>
2026-09-15T07:35:34.118796+08:00 [info] msg: starting_rlog_shard, shard: emqx_acl_sharded
2026-09-15T07:35:34.119886+08:00 [info] msg: Shard fully up, node: ‘emqx@192.168.2.130’, shard: emqx_acl_sharded
2026-09-15T07:35:34.124055+08:00 [info] msg: starting_rlog_shard, shard: emqx_authn_shard
2026-09-15T07:35:34.125268+08:00 [info] msg: Shard fully up, node: ‘emqx@192.168.2.130’, shard: emqx_authn_shard
2026-09-15T07:35:34.134372+08:00 [info] msg: starting_rlog_shard, shard: emqx_telemetry_shard
2026-09-15T07:35:34.136481+08:00 [info] msg: starting_rlog_shard, shard: emqx_psk_shard
2026-09-15T07:35:34.137963+08:00 [info] msg: Shard fully up, node: ‘emqx@192.168.2.130’, shard: emqx_psk_shard
2026-09-15T07:35:34.138150+08:00 [info] msg: emqx_psk_disabled
2026-09-15T07:35:34.139763+08:00 [info] msg: Shard fully up, node: ‘emqx@192.168.2.130’, shard: emqx_telemetry_shard
2026-09-15T08:35:27.108153+08:00 [info] msg: dashboard_login_successful, username: admin
来点日志?
我们的网络状况可能就是短抖动,之前4.4的时候发生过网络分区,导致分区恢复后网关上的数据订阅者收不到,后来升级到5.8.9,虽然发生网络分区后也能恢复,但是也会出现数据断断续续的,就是:假如网关1分钟上送一次数据,那么订阅者这边收到的并不是1分钟一次的数据,而是3 4分钟一次的数据,所以说是断断续续的。。而且,有时候dashboard点击查询某个链接时提示“该客户端不存在”,但是界面确实展示该连接为 已连接。。当踢了该连接或者重启emqx后才恢复正常。
附加一个12号的日志,用户反馈12号0时到8时之间的上送的数据是断断续续的,因为那个时候日志等级是warn,所以可能日志量不多。
2026-09-11T23:16:26.954533+08:00 [warning] msg: cm_registry_node_down, node: ‘emqx@192.168.2.131’
2026-09-11T23:16:26.966967+08:00 [warning] msg: cm_registry_mnesia_down, node: ‘emqx@192.168.2.131’
2026-09-11T23:16:44.921385+08:00 [warning] msg: Stopping mria, reason: heal
2026-09-11T23:16:47.092876+08:00 [critical] msg: Rejoin for autoheal, return: ok, node: ‘emqx@192.168.2.131’
2026-09-11T23:16:47.085479+08:00 [warning] msg: Stopping mria, reason: heal
2026-09-11T23:16:47.914678+08:00 [warning] msg: cm_registry_mnesia_down, node: ‘emqx@192.168.2.132’
2026-09-11T23:16:48.555543+08:00 [critical] msg: Rejoin for autoheal, return: ok, node: ‘emqx@192.168.2.132’
2026-09-11T23:52:18.821035+08:00 [warning] msg: cm_registry_node_down, node: ‘emqx@192.168.2.131’
2026-09-11T23:52:18.835264+08:00 [error] Mnesia(‘emqx@192.168.2.130’): ** ERROR ** mnesia_event got {inconsistent_database, running_partitioned_network, ‘emqx@192.16
8.2.131’}
2026-09-11T23:52:18.835259+08:00 [critical] msg: Core cluster partition, context: running_partitioned_network, from: ‘emqx@192.168.2.131’
2026-09-11T23:52:18.835998+08:00 [warning] msg: alarm_is_activated, message: <<“Partition occurs at node emqx@192.168.2.131”>>, name: partition
2026-09-11T23:52:18.897289+08:00 [warning] msg: cm_registry_mnesia_down, node: ‘emqx@192.168.2.131’
2026-09-11T23:52:36.731242+08:00 [warning] msg: Stopping mria, reason: heal
2026-09-11T23:52:37.686390+08:00 [warning] msg: alarm_is_deactivated, name: partition
2026-09-11T23:52:38.078937+08:00 [critical] msg: Rejoin for autoheal, return: ok, node: ‘emqx@192.168.2.131’
2026-09-12T08:36:01.370860+08:00 [warning] msg: cm_registry_node_down, node: ‘emqx@192.168.2.131’
2026-09-12T08:36:01.620987+08:00 [warning] msg: cm_registry_mnesia_down, node: ‘emqx@192.168.2.131’
2026-09-12T08:36:19.425326+08:00 [warning] msg: Stopping mria, reason: heal
2026-09-12T08:36:21.184718+08:00 [critical] msg: Rejoin for autoheal, return: ok, node: ‘emqx@192.168.2.131’
2026-09-12T08:37:46.053004+08:00 [error] tag: RPC, msg: rpc_channel_error, cause: etimedout, driver: tcp, socket: #Port<0.105>, action: stopping
2026-09-12T09:02:16.776848+08:00 [error] tag: RPC, msg: rpc_channel_error, cause: etimedout, driver: tcp, socket: #Port<0.22523>, action: stopping
2026-09-12T09:02:16.777611+08:00 [warning] msg: cm_registry_node_down, node: ‘emqx@192.168.2.131’
2026-09-12T09:02:16.798733+08:00 [warning] msg: cm_registry_mnesia_down, node: ‘emqx@192.168.2.131’
2026-09-12T09:02:18.284332+08:00 [warning] msg: exclusive_sub_worker_unexpected_info, info: {mnesia_locker,‘emqx@192.168.2.131’,granted}
2026-09-12T09:02:18.284740+08:00 [error] Mnesia(‘emqx@192.168.2.130’): ** ERROR ** mnesia_event got {inconsistent_database, running_partitioned_network, ‘emqx@192.16
8.2.131’}
2026-09-12T09:02:18.284526+08:00 [error] msg: unexpected_info, info: {mnesia_locker,‘emqx@192.168.2.131’,granted}
2026-09-12T09:02:18.284917+08:00 [critical] msg: Core cluster partition, context: running_partitioned_network, from: ‘emqx@192.168.2.131’
2026-09-12T09:02:18.285755+08:00 [warning] msg: alarm_is_activated, message: <<“Partition occurs at node emqx@192.168.2.131”>>, name: partition
2026-09-12T09:02:36.314777+08:00 [warning] msg: Stopping mria, reason: heal
2026-09-12T09:02:37.001740+08:00 [warning] msg: alarm_is_deactivated, name: partition
2026-09-12T09:02:37.899144+08:00 [critical] msg: Rejoin for autoheal, return: ok, node: ‘emqx@192.168.2.131’
2026-09-12T09:02:37.924008+08:00 [warning] msg: Stopping mria, reason: heal
2026-09-12T09:02:38.590421+08:00 [warning] msg: cm_registry_mnesia_down, node: ‘emqx@192.168.2.132’
2026-09-12T09:02:39.331789+08:00 [critical] msg: Rejoin for autoheal, return: ok, node: ‘emqx@192.168.2.132’
2026-09-12T10:13:58.037136+08:00 [warning] msg: cm_registry_node_down, node: ‘emqx@192.168.2.131’
2026-09-12T10:13:58.208356+08:00 [warning] msg: cm_registry_mnesia_down, node: ‘emqx@192.168.2.131’
2026-09-12T10:13:58.859660+08:00 [error] Mnesia(‘emqx@192.168.2.130’): ** ERROR ** mnesia_event got {inconsistent_database, running_partitioned_network, ‘emqx@192.16
8.2.131’}
2026-09-12T10:13:58.859935+08:00 [critical] msg: Core cluster partition, context: running_partitioned_network, from: ‘emqx@192.168.2.131’
2026-09-12T10:13:58.860963+08:00 [warning] msg: alarm_is_activated, message: <<“Partition occurs at node emqx@192.168.2.131”>>, name: partition
2026-09-12T10:14:17.066518+08:00 [warning] msg: Stopping mria, reason: heal
2026-09-12T10:14:17.821254+08:00 [warning] msg: alarm_is_deactivated, name: partition
2026-09-12T10:14:19.176320+08:00 [critical] msg: Rejoin for autoheal, return: ok, node: ‘emqx@192.168.2.131’
这些日志可以确定集群网络分区,并且 07:35:30 触发了 .130 上的 Mria 自动愈合;算是网络短抖动,但还不能直接推断是 5.8.9 的bug。
关键时间线:
- 07:35:11
.131被判定 down,随后 Mnesia down;07:35:12 节点又 up,同时出现running_partitioned_network、inconsistent_database和 partition alarm。 - 07:35:30
.130记录Mria is restarting to join the cluster、stopping_emqx_apps,1883/8883 监听器随之停止;07:35:32–07:35:34 Mria、各分片和监听器重新启动。 - 其他时段也反复出现
.131down、rpc_channel_error ... etimedout、heal/rejoin。客户端的keepalive_timeout/tcp_closed更像网络异常下的结果,不是节点重启原因本身。
因此,3–4 分钟才收到一次数据、Dashboard 显示“客户端不存在”但列表仍是已连接,都可能是分区期间跨节点的会话/订阅/节点状态暂时不一致;要确认消息是否丢失,还需要区分客户端实际连接节点、订阅节点和发布路径。
先不要清空 Mnesia 或关闭自动愈合,
- 说明实际部署方式(物理机、Docker 还是 K8s)、3 个
node.name,以及负载均衡/网关拓扑。 - 提供 3 个节点在同一时间窗(例如
07:35:00–07:36:00)的完整日志,尤其是.131本机是否有容器/OS 重启记录,以及.130触发 heal 前后的日志。 - 在故障前后分别执行
emqx ctl cluster status,贴出running_nodes和stopped_nodes;同时核对节点间集群端口。物理机至少检查 Erlang 分布式端口4370和集群 RPC5370,Docker 的集群 RPC 默认是5369;多节点同机部署还要按节点名后缀检查对应端口。 - 把一条“数据延迟”的完整链路补出来:网关 clientid、发布主题、订阅者 clientid、两者分别连接的节点、QoS/是否持久会话,以及延迟期间各节点的发布/投递日志。
Dashboard 报“客户端不存在”时,也请记录访问的是哪个节点和目标 clientid。
如果能确认确实是 EMQX 进程被外部重启,再补容器编排/系统层事件;目前日志显示的是 .130 因网络分区进入 Mria heal 后主动停止并恢复 EMQX 应用,不是崩溃的 crash dump。
一、部署方式:docker集群部署,docker-compose文件为:
version: '3'
##############################################################################
# EMQX 集群编排文件(host 网络模式)
#
# 注意:
# - 本文件为独立新增文件,不改动原 docker-compose.yml
# - 必须通过 ./install_emqx_cluster.sh 启动(集群配置在 cluster.conf,
# 脚本会导出 ${HOST_IP} / ${EMQX_CLUSTER_SEEDS} 等变量供本文件插值)
# - host 模式:容器直接使用宿主机网络,无端口映射,无需暴露端口
# - EMQX 5.x:集群配置通过环境变量注入(双下划线 __ 映射 HOCON 层级),
# 优先级高于 etc/emqx.conf,无需修改配置文件
##############################################################################
services:
emqx:
image: ${EMQX_IMAGE:-emqx/emqx:5.8.9}
container_name: emqx
network_mode: host
user: root #解决挂载目录权限问题
restart: always
# EMQX 5.x(Erlang/OTP 26)创建线程依赖 clone3 等系统调用,
# Docker 默认 seccomp profile 会拦截,导致启动报错:
# "Failed to create thread: Operation not permitted (1)"
# 解法:仅对本容器放宽 seccomp(EMQX 官方推荐,非 privileged,不影响宿主机安全边界)
security_opt:
- seccomp:unconfined
environment:
# 节点名:emqx@<本机IP>(5.x 用 EMQX_NODE__NAME 双下划线;
# 若误用 EMQX_NODE_NAME 单下划线,节点名回退为 emqx@<hostname>,
# hostname 无法解析将导致 net_kernel 启动失败)
- EMQX_NODE__NAME=emqx@${HOST_IP:?请先执行 install_emqx_cluster.sh}
# Cookie:三台节点必须一致(5.x 用 EMQX_NODE__COOKIE,双下划线)
- EMQX_NODE__COOKIE=${EMQX_COOKIE:?请先执行 install_emqx_cluster.sh}
# 集群名称:三台一致
- EMQX_CLUSTER__NAME=emqxcl
# 集群发现方式:static 静态节点列表(5.x 配置键为 discovery_strategy)
- EMQX_CLUSTER__DISCOVERY_STRATEGY=static
# 集群种子节点列表
- EMQX_CLUSTER__STATIC__SEEDS=${EMQX_CLUSTER_SEEDS:?请先执行 install_emqx_cluster.sh}
- EMQX_LOG__TO=both
volumes:
- $PWD/middle/emqx/data:/opt/emqx/data
- $PWD/middle/emqx/log:/opt/emqx/log
- $PWD/middle/emqx/conf:/opt/emqx/etc
- /etc/localtime:/etc/localtime:ro # 让容器的时钟与宿主机时钟同步,避免时间的问题,ro是read only的意思,就是只读。
healthcheck:
test: ["CMD", "/opt/emqx/bin/emqx", "ctl", "status"]
interval: 5s
timeout: 25s
retries: 5
最外层通过nginx负载均衡代理:
upstream mqtt_servers {
least_conn;
server 192.168.2.130:1883 max_fails=2 fail_timeout=10s;
server 192.168.2.131:1883 max_fails=2 fail_timeout=10s;
server 192.168.2.132:1883 max_fails=2 fail_timeout=10s;
}
server {
listen 18458;
proxy_pass mqtt_servers;
# 启用此项时,对应后端监听器也需要启用 proxy_protocol
proxy_protocol off;
proxy_connect_timeout 10s;
# 默认心跳时间为 10 分钟
proxy_timeout 1800s;
proxy_buffer_size 3M;
tcp_nodelay on;
}
三个node.name: emqx@192.168.2.130、emqx@192.168.2.131、emqx@192.168.2.132
二、 3 个节点在同一时间窗(例如 07:35:00–07:36:00 )的完整日志
131:
2026-09-15T06:03:49.867509+08:00 [info] clientid: universal_ies832.JNXKY_GateWay_001|securemode=1,authType=1,signmethod=HmacSH
A256,timestamp=1756434977086,tenantId=1919995136244723714|, msg: terminate, peername: 192.168.2.130:41082, username: JNXKY_Gat
eWay_001&universal_ies832, reason: {shutdown,tcp_closed}
2026-09-15T07:34:29.535210+08:00 [info] clientid: universal_ies832.JNXKY_GateWay_001|securemode=1,authType=1,signmethod=HmacSH
A256,timestamp=1756434977086,tenantId=1919995136244723714|, msg: terminate, peername: 192.168.2.130:60826, username: JNXKY_Gat
eWay_001&universal_ies832, reason: {shutdown,tcp_closed}
2026-09-15T07:35:11.631275+08:00 [info] Mria(Membership): Node emqx@192.168.2.130 down
2026-09-15T07:35:11.631570+08:00 [warning] msg: cm_registry_node_down, node: 'emqx@192.168.2.130'
2026-09-15T07:35:11.639301+08:00 [info] Mria(Membership): Mnesia emqx@192.168.2.130 down
2026-09-15T07:35:11.648502+08:00 [warning] msg: cm_registry_mnesia_down, node: 'emqx@192.168.2.130'
2026-09-15T07:35:12.665037+08:00 [info] Mria(Membership): Node emqx@192.168.2.130 up
2026-09-15T07:35:12.665642+08:00 [error] msg: unexpected_info, info: {mnesia_locker,'emqx@192.168.2.130',granted}
2026-09-15T07:35:12.665524+08:00 [warning] msg: exclusive_sub_worker_unexpected_info, info: {mnesia_locker,'emqx@192.168.2.130
',granted}
2026-09-15T07:35:12.666789+08:00 [error] Mnesia('emqx@192.168.2.131'): ** ERROR ** mnesia_event got {inconsistent_database, ru
nning_partitioned_network, 'emqx@192.168.2.130'}
2026-09-15T07:35:12.666858+08:00 [critical] msg: Core cluster partition, context: running_partitioned_network, from: 'emqx@192
.168.2.130'
2026-09-15T07:35:12.667800+08:00 [warning] msg: alarm_is_activated, message: <<"Partition occurs at node emqx@192.168.2.130">>
, name: partition
2026-09-15T07:35:14.634557+08:00 [info] msg: mria_node_monitor_confirm, status: up, target_node: 'emqx@192.168.2.130'
2026-09-15T07:35:30.675384+08:00 [info] Mria(Membership): Node emqx@192.168.2.130 healing
2026-09-15T07:35:32.368222+08:00 [info] Mria(Membership): Mnesia emqx@192.168.2.130 up
2026-09-15T07:35:32.368945+08:00 [warning] msg: alarm_is_deactivated, name: partition
2026-09-15T08:29:04.685670+08:00 [info] clientid: universal_ies832.JNXKY_GateWay_001|securemode=1,authType=1,signmethod=HmacSH
A256,timestamp=1756434977086,tenantId=1919995136244723714|, msg: terminate, peername: 192.168.2.130:64474, username: JNXKY_Gat
eWay_001&universal_ies832, reason: {shutdown,keepalive_timeout}
130:
2026-09-15T04:07:46.730471+08:00 [info] clientid: universal_ies832.NBDL_DQ_A1_Gateway_002|securemode=1,authType=1,signmethod=H
macSHA256,timestamp=1755150868383,tenantId=1919995136244723714|, msg: terminate, peername: 192.168.2.130:63676, username: NBDL
_DQ_A1_Gateway_002&universal_ies832, reason: {shutdown,tcp_closed}
2026-09-15T07:35:11.981518+08:00 [info] Mria(Membership): Node emqx@192.168.2.131 down
2026-09-15T07:35:11.982054+08:00 [warning] msg: cm_registry_node_down, node: 'emqx@192.168.2.131'
2026-09-15T07:35:12.046948+08:00 [info] Mria(Membership): Mnesia emqx@192.168.2.131 down
2026-09-15T07:35:12.059106+08:00 [warning] msg: cm_registry_mnesia_down, node: 'emqx@192.168.2.131'
2026-09-15T07:35:12.708508+08:00 [info] Mria(Membership): Node emqx@192.168.2.131 up
2026-09-15T07:35:12.709836+08:00 [warning] msg: exclusive_sub_worker_unexpected_info, info: {mnesia_locker,'emqx@192.168.2.131
',granted}
2026-09-15T07:35:12.710022+08:00 [error] msg: unexpected_info, info: {mnesia_locker,'emqx@192.168.2.131',granted}
2026-09-15T07:35:12.710677+08:00 [critical] msg: Core cluster partition, context: running_partitioned_network, from: 'emqx@192
.168.2.131'
2026-09-15T07:35:12.710667+08:00 [error] Mnesia('emqx@192.168.2.130'): ** ERROR ** mnesia_event got {inconsistent_database, ru
nning_partitioned_network, 'emqx@192.168.2.131'}
2026-09-15T07:35:12.711090+08:00 [warning] msg: alarm_is_activated, message: <<"Partition occurs at node emqx@192.168.2.131">>
, name: partition
2026-09-15T07:35:14.985347+08:00 [info] msg: mria_node_monitor_confirm, status: up, target_node: 'emqx@192.168.2.131'
2026-09-15T07:35:30.718879+08:00 [notice] msg: Mria is restarting to join the cluster, seed: 'emqx@192.168.2.132'
2026-09-15T07:35:30.719792+08:00 [warning] msg: Stopping mria, reason: heal
2026-09-15T07:35:30.720065+08:00 [notice] msg: stopping_emqx_apps
2026-09-15T07:35:30.766199+08:00 [notice] Application: quicer. Exited: stopped. Type: temporary.
2026-09-15T07:35:30.772166+08:00 [notice] Application: emqx_connector_jwt. Exited: stopped. Type: temporary.
2026-09-15T07:35:30.779391+08:00 [notice] Application: emqx_psk. Exited: stopped. Type: temporary.
2026-09-15T07:35:30.785414+08:00 [notice] Application: emqx_telemetry. Exited: stopped. Type: temporary.
2026-09-15T07:35:30.785945+08:00 [notice] Application: emqx_gateway_mqttsn. Exited: stopped. Type: temporary.
2026-09-15T07:35:30.786283+08:00 [notice] Application: emqx_gateway_lwm2m. Exited: stopped. Type: temporary.
2026-09-15T07:35:30.786457+08:00 [notice] Application: emqx_gateway_coap. Exited: stopped. Type: temporary.
2026-09-15T07:35:30.786614+08:00 [notice] Application: emqx_gateway_stomp. Exited: stopped. Type: temporary.
2026-09-15T07:35:30.786978+08:00 [notice] Application: emqx_gateway_exproto. Exited: stopped. Type: temporary.
2026-09-15T07:35:30.794057+08:00 [notice] Application: emqx_gateway. Exited: stopped. Type: temporary.
2026-09-15T07:35:30.799614+08:00 [notice] Application: emqx_auth_mysql. Exited: stopped. Type: temporary.
2026-09-15T07:35:30.799976+08:00 [notice] Application: emqx_mysql. Exited: stopped. Type: temporary.
2026-09-15T07:35:30.805934+08:00 [notice] Application: emqx_prometheus. Exited: stopped. Type: temporary.
2026-09-15T07:35:30.811385+08:00 [notice] Application: emqx_auth_mnesia. Exited: stopped. Type: temporary.
2026-09-15T07:35:30.817364+08:00 [notice] Application: emqx_auth_redis. Exited: stopped. Type: temporary.
2026-09-15T07:35:30.823712+08:00 [notice] Application: emqx_auth_http. Exited: stopped. Type: temporary.
2026-09-15T07:35:30.830293+08:00 [notice] Application: emqx_auth_mongodb. Exited: stopped. Type: temporary.
2026-09-15T07:35:30.835981+08:00 [notice] Application: emqx_auth_jwt. Exited: stopped. Type: temporary.
2026-09-15T07:35:30.841852+08:00 [notice] Application: emqx_auth_ldap. Exited: stopped. Type: temporary.
2026-09-15T07:35:30.847862+08:00 [notice] Application: emqx_auth_postgresql. Exited: stopped. Type: temporary.
2026-09-15T07:35:30.852310+08:00 [info] msg: stopping_http_connector, connector: <<"emqx_authn_http:64516">>
2026-09-15T07:35:30.871149+08:00 [notice] Application: emqx_auth. Exited: stopped. Type: temporary.
2026-09-15T07:35:30.877198+08:00 [notice] Application: emqx_auto_subscribe. Exited: stopped. Type: temporary.
2026-09-15T07:35:30.886367+08:00 [notice] Application: emqx_dashboard. Exited: stopped. Type: temporary.
2026-09-15T07:35:30.904529+08:00 [notice] Application: emqx_rule_engine. Exited: stopped. Type: temporary.
2026-09-15T07:35:30.919301+08:00 [notice] Application: emqx_bridge. Exited: stopped. Type: temporary.
2026-09-15T07:35:30.920397+08:00 [notice] Application: emqx_ldap. Exited: stopped. Type: temporary.
2026-09-15T07:35:30.927153+08:00 [notice] Application: emqx_connector. Exited: stopped. Type: temporary.
2026-09-15T07:35:30.927721+08:00 [notice] Application: emqx_mongodb. Exited: stopped. Type: temporary.
2026-09-15T07:35:30.928356+08:00 [info] msg: exhook_mgr_terminated, reason: shutdown, servers: #{}
2026-09-15T07:35:30.934275+08:00 [notice] Application: emqx_exhook. Exited: stopped. Type: temporary.
2026-09-15T07:35:30.957399+08:00 [notice] Application: emqx_retainer. Exited: stopped. Type: temporary.
2026-09-15T07:35:30.976163+08:00 [notice] Application: emqx_modules. Exited: stopped. Type: temporary.
2026-09-15T07:35:30.982446+08:00 [notice] Application: emqx_slow_subs. Exited: stopped. Type: temporary.
2026-09-15T07:35:30.988411+08:00 [notice] Application: emqx_management. Exited: stopped. Type: temporary.
2026-09-15T07:35:30.994389+08:00 [notice] Application: emqx_plugins. Exited: stopped. Type: temporary.
2026-09-15T07:35:30.994928+08:00 [notice] Application: emqx_bridge_mqtt. Exited: stopped. Type: temporary.
2026-09-15T07:35:30.996541+08:00 [notice] Application: emqx_bridge_http. Exited: stopped. Type: temporary.
2026-09-15T07:35:30.996830+08:00 [notice] Application: emqx_redis. Exited: stopped. Type: temporary.
2026-09-15T07:35:30.997059+08:00 [notice] Application: emqx_postgresql. Exited: stopped. Type: temporary.
2026-09-15T07:35:31.014404+08:00 [notice] Application: emqx_resource. Exited: stopped. Type: temporary.
2026-09-15T07:35:31.016106+08:00 [notice] tcp:default stopped on 0.0.0.0:1883
2026-09-15T07:35:31.020814+08:00 [notice] ssl:default stopped on 0.0.0.0:8883
2026-09-15T07:35:31.100237+08:00 [notice] Application: emqx. Exited: stopped. Type: permanent.
2026-09-15T07:35:31.101037+08:00 [notice] Application: emqx_ds_backends. Exited: stopped. Type: temporary.
2026-09-15T07:35:31.115741+08:00 [notice] Application: emqx_durable_storage. Exited: stopped. Type: temporary.
2026-09-15T07:35:31.116161+08:00 [notice] Application: emqx_utils. Exited: stopped. Type: temporary.
2026-09-15T07:35:31.126884+08:00 [notice] Application: cowboy. Exited: stopped. Type: temporary.
2026-09-15T07:35:31.140147+08:00 [notice] Application: ranch. Exited: stopped. Type: temporary.
2026-09-15T07:35:31.155944+08:00 [notice] Application: bcrypt. Exited: stopped. Type: temporary.
2026-09-15T07:35:31.156355+08:00 [notice] Application: emqx_http_lib. Exited: stopped. Type: temporary.
2026-09-15T07:35:31.166931+08:00 [notice] Application: gproc. Exited: stopped. Type: temporary.
2026-09-15T07:35:31.181337+08:00 [notice] Application: esockd. Exited: stopped. Type: permanent.
2026-09-15T07:35:31.196590+08:00 [notice] Application: emqx_conf. Exited: stopped. Type: temporary.
2026-09-15T07:35:31.209492+08:00 [notice] Application: ekka. Exited: stopped. Type: temporary.
2026-09-15T07:35:31.253510+08:00 [notice] msg: Mria is stopped
2026-09-15T07:35:31.264093+08:00 [notice] Application: mria. Exited: stopped. Type: temporary.
2026-09-15T07:35:31.295314+08:00 [notice] Application: mnesia. Exited: stopped. Type: temporary.
2026-09-15T07:35:32.418414+08:00 [notice] msg: Starting mria
2026-09-15T07:35:32.420950+08:00 [notice] msg: Starting mnesia
2026-09-15T07:35:32.426000+08:00 [notice] msg: Starting shards
2026-09-15T07:35:32.750320+08:00 [info] msg: Setting RLOG shard config, tables: ['$mria_rlog_sync',mria_schema], shard: '$mria
_meta_shard'
2026-09-15T07:35:32.751141+08:00 [info] msg: Converging schema
2026-09-15T07:35:32.753221+08:00 [info] msg: Setting RLOG shard config, tables: ['$mria_rlog_sync',mria_schema], shard: '$mria
_meta_shard'
2026-09-15T07:35:32.755023+08:00 [info] msg: Setting RLOG shard config, tables: [bpapi,emqx_app,emqx_banned,emqx_banned_rules,
emqx_trace], shard: emqx_common_shard
2026-09-15T07:35:32.756781+08:00 [info] msg: Setting RLOG shard config, tables: [cluster_rpc_commit,cluster_rpc_mfa], shard: e
mqx_cluster_rpc_shard
2026-09-15T07:35:32.758337+08:00 [info] msg: Setting RLOG shard config, tables: [cluster_rpc_commit,cluster_rpc_mfa], shard: e
mqx_cluster_rpc_shard
2026-09-15T07:35:32.759630+08:00 [info] msg: Setting RLOG shard config, tables: [emqx_acl], shard: emqx_acl_sharded
2026-09-15T07:35:32.760858+08:00 [info] msg: Setting RLOG shard config, tables: [emqx_activated_alarm,emqx_dashboard_monitor,e
mqx_deactivated_alarm,emqx_delayed,emqx_ds_builtin_local_metadata_tab,emqx_ds_builtin_local_timestamp_tab], shard: undefined
2026-09-15T07:35:32.762271+08:00 [info] msg: Setting RLOG shard config, tables: [emqx_admin,emqx_admin_jwt], shard: emqx_dashb
oard_shard
2026-09-15T07:35:32.763673+08:00 [info] msg: Setting RLOG shard config, tables: [emqx_admin,emqx_admin_jwt], shard: emqx_dashb
oard_shard
2026-09-15T07:35:32.765308+08:00 [info] msg: Setting RLOG shard config, tables: [bpapi,emqx_app,emqx_banned,emqx_banned_rules,
emqx_trace], shard: emqx_common_shard
2026-09-15T07:35:32.766687+08:00 [info] msg: Setting RLOG shard config, tables: [emqx_authn_mnesia,emqx_authn_scram_mnesia], s
hard: emqx_authn_shard
2026-09-15T07:35:32.768086+08:00 [info] msg: Setting RLOG shard config, tables: [emqx_authn_mnesia,emqx_authn_scram_mnesia], s
hard: emqx_authn_shard
2026-09-15T07:35:32.769597+08:00 [info] msg: Setting RLOG shard config, tables: [bpapi,emqx_app,emqx_banned,emqx_banned_rules,
emqx_trace], shard: emqx_common_shard
2026-09-15T07:35:32.771464+08:00 [info] msg: Setting RLOG shard config, tables: [bpapi,emqx_app,emqx_banned,emqx_banned_rules,
emqx_trace], shard: emqx_common_shard
2026-09-15T07:35:32.772918+08:00 [info] msg: Setting RLOG shard config, tables: [emqx_channel_registry], shard: emqx_cm_shard
2026-09-15T07:35:32.774947+08:00 [info] msg: Setting RLOG shard config, tables: [emqx_activated_alarm,emqx_dashboard_monitor,e
mqx_deactivated_alarm,emqx_delayed,emqx_ds_builtin_local_metadata_tab,emqx_ds_builtin_local_timestamp_tab], shard: undefined
2026-09-15T07:35:32.776438+08:00 [info] msg: Setting RLOG shard config, tables: [emqx_activated_alarm,emqx_dashboard_monitor,e
mqx_deactivated_alarm,emqx_delayed,emqx_ds_builtin_local_metadata_tab,emqx_ds_builtin_local_timestamp_tab], shard: undefined
2026-09-15T07:35:32.777863+08:00 [info] msg: Setting RLOG shard config, tables: [emqx_activated_alarm,emqx_dashboard_monitor,e
mqx_deactivated_alarm,emqx_delayed,emqx_ds_builtin_local_metadata_tab,emqx_ds_builtin_local_timestamp_tab], shard: undefined
2026-09-15T07:35:32.780480+08:00 [info] msg: Setting RLOG shard config, tables: [emqx_activated_alarm,emqx_dashboard_monitor,e
mqx_deactivated_alarm,emqx_delayed,emqx_ds_builtin_local_metadata_tab,emqx_ds_builtin_local_timestamp_tab], shard: undefined
2026-09-15T07:35:32.782118+08:00 [info] msg: Setting RLOG shard config, tables: [emqx_activated_alarm,emqx_dashboard_monitor,e
mqx_deactivated_alarm,emqx_delayed,emqx_ds_builtin_local_metadata_tab,emqx_ds_builtin_local_timestamp_tab], shard: undefined
2026-09-15T07:35:32.783405+08:00 [info] msg: Setting RLOG shard config, tables: [emqx_exclusive_subscription,emqx_exclusive_su
bscription_v2], shard: emqx_exclusive_shard
2026-09-15T07:35:32.784716+08:00 [info] msg: Setting RLOG shard config, tables: [emqx_exclusive_subscription,emqx_exclusive_su
bscription_v2], shard: emqx_exclusive_shard
2026-09-15T07:35:32.786511+08:00 [info] msg: Setting RLOG shard config, tables: [emqx_psk], shard: emqx_psk_shard
2026-09-15T07:35:32.788761+08:00 [info] msg: Setting RLOG shard config, tables: [emqx_retainer_index,emqx_retainer_index_meta,
emqx_retainer_message], shard: emqx_retainer_shard
2026-09-15T07:35:32.790795+08:00 [info] msg: Setting RLOG shard config, tables: [emqx_retainer_index,emqx_retainer_index_meta,
emqx_retainer_message], shard: emqx_retainer_shard
2026-09-15T07:35:32.793113+08:00 [info] msg: Setting RLOG shard config, tables: [emqx_retainer_index,emqx_retainer_index_meta,
emqx_retainer_message], shard: emqx_retainer_shard
2026-09-15T07:35:32.794964+08:00 [info] msg: Setting RLOG shard config, tables: [emqx_route,emqx_route_filters,emqx_routing_no
de,emqx_trie], shard: route_shard
2026-09-15T07:35:32.799455+08:00 [info] msg: Setting RLOG shard config, tables: [emqx_route,emqx_route_filters,emqx_routing_no
de,emqx_trie], shard: route_shard
2026-09-15T07:35:32.804643+08:00 [info] msg: Setting RLOG shard config, tables: [emqx_route,emqx_route_filters,emqx_routing_no
de,emqx_trie], shard: route_shard
2026-09-15T07:35:32.806478+08:00 [info] msg: Setting RLOG shard config, tables: [emqx_shared_subscription], shard: emqx_shared
_sub_shard
2026-09-15T07:35:32.808100+08:00 [info] msg: Setting RLOG shard config, tables: [emqx_telemetry], shard: emqx_telemetry_shard
2026-09-15T07:35:32.810504+08:00 [info] msg: Setting RLOG shard config, tables: [bpapi,emqx_app,emqx_banned,emqx_banned_rules,
emqx_trace], shard: emqx_common_shard
2026-09-15T07:35:32.812222+08:00 [info] msg: Setting RLOG shard config, tables: [emqx_route,emqx_route_filters,emqx_routing_no
de,emqx_trie], shard: route_shard
2026-09-15T07:35:32.813465+08:00 [info] Mria(Membership): Node emqx@192.168.2.132 up
2026-09-15T07:35:32.813902+08:00 [info] msg: starting_rlog_shard, shard: '$mria_meta_shard'
2026-09-15T07:35:32.814677+08:00 [info] Mria(Membership): Node emqx@192.168.2.131 up
2026-09-15T07:35:32.815562+08:00 [notice] msg: Mria is running
2026-09-15T07:35:32.815590+08:00 [info] msg: Shard fully up, node: 'emqx@192.168.2.130', shard: '$mria_meta_shard'
2026-09-15T07:35:32.816013+08:00 [info] msg: Starting ekka
2026-09-15T07:35:32.816932+08:00 [notice] msg: Mria has joined the cluster, status: #{members => [{member,'emqx@192.168.2.130'
,undefined,<<0,6,91,121,227,228,32,37,156,25,0,11,150,197,0,0>>,3325906226,up,running,{1789,428932,812889},core},{member,'emqx
@192.168.2.132',undefined,undefined,undefined,up,running,{1789,428932,814655},core}],running_nodes => ['emqx@192.168.2.130','e
mqx@192.168.2.131','emqx@192.168.2.132'],rlog => #{role => core,backend => rlog,imbalance => []},partitions => [],rebalance_st
atus => not_started,stopped_nodes => []}, seed: 'emqx@192.168.2.132'
2026-09-15T07:35:32.817923+08:00 [info] msg: Ekka is running
2026-09-15T07:35:32.818251+08:00 [notice] msg: (re)starting_emqx_apps
2026-09-15T07:35:32.830257+08:00 [info] msg: starting_rlog_shard, shard: emqx_cluster_rpc_shard
2026-09-15T07:35:32.831438+08:00 [info] msg: wait_for_cluster_rpc_shard, result: ok
2026-09-15T07:35:32.831530+08:00 [info] msg: Shard fully up, node: 'emqx@192.168.2.130', shard: emqx_cluster_rpc_shard
2026-09-15T07:35:32.831783+08:00 [info] msg: wait_for_cluster_rpc_tables, result: ok
2026-09-15T07:35:32.844566+08:00 [info] msg: sync_cluster_conf_success, data_dir: data, has_deprecated_file: false, tnx_id: 12
, local_release: v5.8.9, remote_release: v5.8.9, synced_from_node: 'emqx@192.168.2.131'
2026-09-15T07:35:33.189582+08:00 [info] msg: starting_rlog_shard, shard: emqx_common_shard
2026-09-15T07:35:33.190482+08:00 [info] msg: Shard fully up, node: 'emqx@192.168.2.130', shard: emqx_common_shard
2026-09-15T07:35:33.245725+08:00 [info] msg: starting_rlog_shard, shard: route_shard
2026-09-15T07:35:33.247314+08:00 [info] msg: Shard fully up, node: 'emqx@192.168.2.130', shard: route_shard
2026-09-15T07:35:33.248274+08:00 [info] msg: routing_schema_used, schema: v2
2026-09-15T07:35:33.265367+08:00 [info] msg: starting_rlog_shard, shard: emqx_shared_sub_shard
2026-09-15T07:35:33.266478+08:00 [info] msg: Shard fully up, node: 'emqx@192.168.2.130', shard: emqx_shared_sub_shard
2026-09-15T07:35:33.304639+08:00 [info] msg: starting_rlog_shard, shard: emqx_exclusive_shard
2026-09-15T07:35:33.305567+08:00 [info] msg: Shard fully up, node: 'emqx@192.168.2.130', shard: emqx_exclusive_shard
2026-09-15T07:35:33.317769+08:00 [info] msg: starting_rlog_shard, shard: emqx_cm_shard
2026-09-15T07:35:33.318900+08:00 [info] msg: Shard fully up, node: 'emqx@192.168.2.130', shard: emqx_cm_shard
2026-09-15T07:35:33.362515+08:00 [info] msg: wait_for_cluster_rpc_shard, result: ok
2026-09-15T07:35:33.363060+08:00 [info] msg: wait_for_cluster_rpc_tables, result: ok
2026-09-15T07:35:33.364225+08:00 [info] msg: CMD_overridden, cmd: observer, mf: {emqx_observer_cli,cmd}
2026-09-15T07:35:33.364660+08:00 [info] msg: CMD_overridden, cmd: cluster_call, mf: {emqx_conf_cli,admins}
2026-09-15T07:35:33.365011+08:00 [info] msg: CMD_overridden, cmd: conf, mf: {emqx_conf_cli,conf}
2026-09-15T07:35:33.373168+08:00 [info] msg: starting_rlog_shard, shard: emqx_retainer_shard
2026-09-15T07:35:33.374862+08:00 [info] msg: Shard fully up, node: 'emqx@192.168.2.130', shard: emqx_retainer_shard
2026-09-15T07:35:33.403933+08:00 [info] msg: starting_rlog_shard, shard: emqx_dashboard_shard
2026-09-15T07:35:33.405299+08:00 [info] msg: Shard fully up, node: 'emqx@192.168.2.130', shard: emqx_dashboard_shard
2026-09-15T07:35:33.408466+08:00 [info] msg: loading_desc, file: /opt/emqx/lib/emqx_dashboard-5.2.3/priv/desc.en.hocon
2026-09-15T07:35:33.632521+08:00 [info] msg: loading_desc, file: /opt/emqx/lib/emqx_dashboard-5.2.3/priv/desc.zh.hocon
2026-09-15T07:35:33.849544+08:00 [info] msg: started_listener_ok, name: 'http:dashboard', pid: <0.760398.0>, port: 18083
2026-09-15T07:35:33.992616+08:00 [info] msg: regenerate_dispatch, listeners: ['http:dashboard'], i18n_lang: en, elapsed_ms: 14
2
2026-09-15T07:35:34.008294+08:00 [info] msg: starting_http_connector, config: #{ssl => #{depth => 10,verify => verify_peer,hib
ernate_after => 5000,enable => false,ciphers => [],log_level => notice,versions => ['tlsv1.3','tlsv1.2'],secure_renegotiate =>
true,reuse_sessions => true},connect_timeout => 15000,mechanism => password_based,pool_size => 8,enable => true,body => #{pas
sword => <<"${password}">>,username => <<"${username}">>,clientid => <<"${clientid}">>,peerhost => <<"${peerhost}">>},headers
=> #{<<"content-type">> => <<"application/json">>},url => <<"http://192.168.2.130:9001/v1/mqtt/auth">>,method => post,backend
=> http,request_timeout => 5000,request_base => #{port => 9001,scheme => http,host => {192,168,2,130}},pool_type => random,ena
ble_pipelining => 100}, connector: <<"emqx_authn_http:125427">>
2026-09-15T07:35:34.118796+08:00 [info] msg: starting_rlog_shard, shard: emqx_acl_sharded
2026-09-15T07:35:34.119886+08:00 [info] msg: Shard fully up, node: 'emqx@192.168.2.130', shard: emqx_acl_sharded
2026-09-15T07:35:34.124055+08:00 [info] msg: starting_rlog_shard, shard: emqx_authn_shard
2026-09-15T07:35:34.125268+08:00 [info] msg: Shard fully up, node: 'emqx@192.168.2.130', shard: emqx_authn_shard
2026-09-15T07:35:34.134372+08:00 [info] msg: starting_rlog_shard, shard: emqx_telemetry_shard
2026-09-15T07:35:34.136481+08:00 [info] msg: starting_rlog_shard, shard: emqx_psk_shard
2026-09-15T07:35:34.137963+08:00 [info] msg: Shard fully up, node: 'emqx@192.168.2.130', shard: emqx_psk_shard
2026-09-15T07:35:34.138150+08:00 [info] msg: emqx_psk_disabled
2026-09-15T07:35:34.139763+08:00 [info] msg: Shard fully up, node: 'emqx@192.168.2.130', shard: emqx_telemetry_shard
2026-09-15T08:35:27.108153+08:00 [info] msg: dashboard_login_successful, username: admin
132:
2026-09-14T23:24:34.885293+08:00 [info] clientid: universal_ies832.SZMXSGYY_008_Gateway|securemode=1,authType=1,signmethod=Hma
cSHA256,timestamp=1757991613407,tenantId=1919995136244723714|, msg: terminate, peername: 192.168.2.130:48782, username: SZMXSG
YY_008_Gateway&universal_ies832, reason: {shutdown,keepalive_timeout}
2026-09-15T07:35:14.665508+08:00 [info] msg: mria_monitor_suspect, target_node: 'emqx@192.168.2.130', from_node: 'emqx@192.168
.2.131'
2026-09-15T07:35:14.971468+08:00 [info] msg: mria_monitor_suspect, target_node: 'emqx@192.168.2.131', from_node: 'emqx@192.168
.2.130'
2026-09-15T07:35:15.701034+08:00 [info] msg: mria_autoheal_report_partition, node: 'emqx@192.168.2.130'
2026-09-15T07:35:15.702027+08:00 [info] msg: mria_autoheal_report_partition, node: 'emqx@192.168.2.131'
2026-09-15T07:35:30.704136+08:00 [info] msg: mria_autoheal_partition, cliques: [['emqx@192.168.2.131','emqx@192.168.2.132'],['
emqx@192.168.2.130','emqx@192.168.2.132']]
2026-09-15T07:35:30.704578+08:00 [info] Mria(Autoheal): Healing partition: [['emqx@192.168.2.131','emqx@192.168.2.132'],['emqx
@192.168.2.130','emqx@192.168.2.132']]
2026-09-15T07:35:30.704820+08:00 [info] msg: Rebooting minority, nodes: ['emqx@192.168.2.130','emqx@192.168.2.132']
2026-09-15T07:35:30.718879+08:00 [notice] msg: Mria is restarting to join the cluster, seed: 'emqx@192.168.2.132'
2026-09-15T07:35:30.708435+08:00 [info] Mria(Membership): Node emqx@192.168.2.130 healing
2026-09-15T07:35:30.719792+08:00 [warning] msg: Stopping mria, reason: heal
2026-09-15T07:35:30.720065+08:00 [notice] msg: stopping_emqx_apps
2026-09-15T07:35:31.279774+08:00 [info] Mria(Membership): Mnesia emqx@192.168.2.130 down
2026-09-15T07:35:31.280218+08:00 [warning] msg: cm_registry_mnesia_down, node: 'emqx@192.168.2.130'
2026-09-15T07:35:32.376188+08:00 [info] Mria(Membership): Mnesia emqx@192.168.2.130 up
2026-09-15T07:35:32.816932+08:00 [notice] msg: Mria has joined the cluster, status: #{members => [{member,'emqx@192.168.2.130'
,undefined,<<0,6,91,121,227,228,32,37,156,25,0,11,150,197,0,0>>,3325906226,up,running,{1789,428932,812889},core},{member,'emqx
@192.168.2.132',undefined,undefined,undefined,up,running,{1789,428932,814655},core}],running_nodes => ['emqx@192.168.2.130','e
mqx@192.168.2.131','emqx@192.168.2.132'],rlog => #{role => core,backend => rlog,imbalance => []},partitions => [],rebalance_st
atus => not_started,stopped_nodes => []}, seed: 'emqx@192.168.2.132'
2026-09-15T07:35:32.807007+08:00 [critical] msg: Rejoin for autoheal, return: ok, node: 'emqx@192.168.2.130'
2026-09-15T07:35:32.807320+08:00 [critical] msg: Rejoin for autoheal, return: ignore, node: 'emqx@192.168.2.132'
131:没有手动容器/OS重启记录:
docker ps -a | grep emqx
92cfb8e53b16 emqx/emqx:5.8.9 “/usr/bin/docker-ent…” 7 days ago Up 3 days (healthy) emqx
三、docker集群rpc-5369监听:
lsof -i:5369
COMMAND PID USER FD TYPE DEVICE SIZE/OFF NODE NAME
beam.smp 773212 root 52u IPv4 2486178862 0t0 TCP *:5369 (LISTEN)
beam.smp 773212 root 69u IPv4 2493473252 0t0 TCP TDengine02:34748->TDengine03:5369 (ESTABLISHED)
beam.smp 773212 root 70u IPv4 2499298626 0t0 TCP TDengine02:15144->TDengine03:5369 (ESTABLISHED)
beam.smp 773212 root 72u IPv4 2497054696 0t0 TCP TDengine02:5369->TDengine03:3422 (ESTABLISHED)
beam.smp 773212 root 73u IPv4 2503231545 0t0 TCP TDengine02:5369->TDengine01:59320 (ESTABLISHED)
beam.smp 773212 root 78u IPv4 2502242863 0t0 TCP TDengine02:38220->TDengine01:5369 (ESTABLISHED)
beam.smp 773212 root 80u IPv4 2493295070 0t0 TCP TDengine02:5369->TDengine01:48878 (ESTABLISHED)
beam.smp 773212 root 86u IPv4 2516258536 0t0 TCP TDengine02:63064->TDengine01:5369 (ESTABLISHED)
beam.smp 773212 root 90u IPv4 2488617909 0t0 TCP TDengine02:51358->TDengine01:5369 (ESTABLISHED)
beam.smp 773212 root 94u IPv4 2502242887 0t0 TCP TDengine02:18036->TDengine01:5369 (ESTABLISHED)
beam.smp 773212 root 95u IPv4 2503378716 0t0 TCP TDengine02:49172->TDengine01:5369 (ESTABLISHED)
beam.smp 773212 root 101u IPv4 2516546674 0t0 TCP TDengine02:29998->TDengine01:5369 (ESTABLISHED)
beam.smp 773212 root 103u IPv4 2486182276 0t0 TCP TDengine02:13220->TDengine01:5369 (ESTABLISHED)
beam.smp 773212 root 104u IPv4 2486182289 0t0 TCP TDengine02:13222->TDengine01:5369 (ESTABLISHED)
beam.smp 773212 root 105u IPv4 2486186231 0t0 TCP TDengine02:5369->TDengine03:41286 (ESTABLISHED)
beam.smp 773212 root 106u IPv4 2486183178 0t0 TCP TDengine02:36954->TDengine01:5369 (ESTABLISHED)
beam.smp 773212 root 114u IPv4 2493464292 0t0 TCP TDengine02:58156->TDengine01:5369 (ESTABLISHED)
emqx ctl cluster status
Cluster status: #{running_nodes =>
['emqx@192.168.2.130','emqx@192.168.2.131',
'emqx@192.168.2.132'],
stopped_nodes => []}
以及延迟期间各节点的发布/投递日志 这个要怎么查? 目前看暂时没有断断续续的了,要想复现还得等后面啥时候网络抖动
先用集群级 Trace 抓,不需要要把三个节点的全局日志都开到 debug。你们现在不需要改配置,下一次出现“3–4 分钟才收到一次”时,提前开一个短时窗口按下面采集。
- 先确定完整主题和两端 Client ID:网关发布端、订阅端;同时记下两端各自连接的 node.name、QoS、订阅过滤器,以及 MQTT 版本、clean_start 和 session_expiry。
- 优先用主题做集群级追踪;如果主题还不能确定,就分别追踪网关和订阅者 Client ID。
在 Dashboard 操作:进入“问题分析 → 日志追踪”,创建 Topic 或 Client ID 追踪,起止时间覆盖复现窗口;集群环境下载时要分别选择三个节点的日志。
3. 复现时让发布端每条消息带唯一序号和发送时间,例如 seq=20260915-001、sent_at=…;订阅端记录实际收到的 seq 和 receive_at。不要只看“客户端是否收到”,否则无法判断消息是在发布前、集群转发中,还是到达订阅连接后延迟。
- 出现一次延迟后
然后在 Dashboard 的“问题分析 → 日志追踪”里查看或下载三个节点的 Trace;不要只下载当前 Dashboard 节点。Docker 也可以从容器实际数据目录下的 data/trace 取文件。
重点按同一个 seq 对齐三个节点的时间线:
- Trace 里没有网关的 publish:先查网关连接、发布端 TCP 链路或网关本身。
- 网关节点有 publish,但没有跨节点 forward/deliver:重点查订阅关系、节点间 RPC 和分区期间的路由状态。
- 已出现 forward/deliver,但订阅端 Trace 或客户端日志晚几分钟:重点查订阅端连接节点、会话队列、QoS ACK,以及客户端到该节点的网络。
- 同时把三个节点对应时间窗的 EMQX 日志和
emqx ctl cluster status --json一起保留,特别关注node_down、running_partitioned_network、mria_autoheal、rpc_channel_error和keepalive_timeout。
你这次还没有复现,不用为了“等网络抖动”长期开 Trace;按上面的 900 秒窗口抓到一条完整链路,就能判断延迟发生在哪一跳。文档:日志追踪 和 CLI Trace。
好的,感谢大佬!
由于5.8.9是最后一个开源版本,而且我们的连接数不多,所以暂时还没有升级企业版的计划。我们计划先找网络部门相关同事看看能不能解决网络短抖动的问题,如果还不行,计划改成单机版试试。
嗯嗯,如果连接数不多,推荐单机版本哈。
