[
https://issues.jboss.org/browse/JGRP-1902?page=com.atlassian.jira.plugin....
]
Dan Berindei commented on JGRP-1902:
------------------------------------
{{timeout}} doesn't work like that in {{FD_ALL}} and {{FD_ALL2}}. {{FD_ALL}} checks
the timeout every {{timeout_check_interval}} millis, so the maximum failure detection time
is {{timeout + timeout_check_interval}} millis. {{FD_ALL2}} only checks the timeout every
{{timeout}} millis, so the maximum detection time is {{2 * timeout}} millis.
I believe messages getting lost only affects the _minimum_ failure detection time (and the
chance of false positives).
Simplify failure detection and merge timeout configuration
----------------------------------------------------------
Key: JGRP-1902
URL:
https://issues.jboss.org/browse/JGRP-1902
Project: JGroups
Issue Type: Enhancement
Affects Versions: 3.6
Reporter: Dan Berindei
Assignee: Bela Ban
Priority: Minor
Fix For: 3.6.2
FD/FD_ALL/FD_ALL2/FD_SOCK javadoc doesn't give any guidance as to how long it would
take to detect a leaving member. MERGE2/MERGE3 javadoc also doesn't say how much it
would take to detect that the network has healed.
For an example of how misleading the current settings can be, I have seen MERGE3 take
more than 20s to merge two partitions with min_interval=1000 and max_interval=5000. FD
also detects a leaver after {{timeout * max_tries}} in the best case, and twice that if 2
consecutive nodes (in the members list) leave at the same time.
The maximum time it takes to detect a leaver is of particular interest to Infinispan
users, because Infinispan is supposed to protect against nodes leaving. But if the users
don't configure a high enough RPC timeout in Infinispan, we don't get to detect
the node leaving.
Ideally, the user should be able to specify a maximum detection time, and the protocol
should adjust the existing settings to meet that (most of the time).
--
This message was sent by Atlassian JIRA
(v6.3.11#6341)