| Commit message (Collapse) | Author | Age | Files | Lines |
| |
|
|
|
|
| |
This only happens when using basic.get but would crash the quorum queue
when the delivery limit was reached due to the transient basic.get
consumer being removed
|
| |\
| |
| | |
Introduce rabbit_upgrade_preparation
|
| | |
| |
| |
| | |
Part of rabbitmq/rabbitmq-cli#408.
|
| | | |
|
| | |
| |
| |
| |
| |
| |
| |
| |
| |
| | |
Rather than wait a fixed 2000ms, poll the test condition for up to
5000ms.
Also switch from a raw message send in rabbit.erl to a gen_event:call/4
to the lager backend. I had hoped this would behave synchronously, which
it does not appear to, but at least we now get a value back from the
call.
|
| |\ \
| | |
| | | |
Use new `inet_tcp_proxy_dist`
|
| | | |
| | |
| | |
| | |
| | |
| | |
| | |
| | |
| | |
| | |
| | | |
`mark_as_enabled_remotely()`
The rounded value of `(timer:now_diff(T1, T0) div 1000)` was often 0,
leading to no decrease in timeout, i.e. the equivalent of an infinite
loop. Now, the division happens after the substraction.
Also, to avoid to much hammering, we sleep for one second between
retries.
|
| |/ /
| |
| |
| | |
See rabbitmq/inet_tcp_proxy#3 and rabbitmq/rabbitmq-ct-helpers#39.
|
| |/ |
|
| |\
| |
| | |
Fix schema sync for parameters
|
| |/
|
|
|
|
| |
Fixes the issue described here:
https://github.com/rabbitmq/rabbitmq-schema-definition-sync/issues/2#issuecomment-613075591
|
| |
|
|
|
|
|
|
|
|
|
| |
When asserting that connections to a down virtual host cannot be open
we compete against a desired behavior: quick recovery of the virtual host
process tree.
Unlike with the "X eventually happens or we fail with a timeout" kind of scenarios,
here the test must catch a very specific time window. Instead of exposing
virtual host internals to the test so that it's more observable from
the outside we concluded that these tests are not critically important to have.
|
| | |
|
| |
|
|
|
| |
and is quite coarse grained, meaning it is hard to keep track of
node state for its duration.
|
| |\
| |
| | |
Remove flaky tests
|
| | |
| |
| |
| |
| |
| |
| |
| |
| |
| |
| |
| |
| |
| |
| |
| | |
After looking at the last 20 builds in the server-release:v3.9.x
pipeline, 2 builds succeeded on the first run. Even though we retry
each build 3 times, out of the last 50 builds, 28 passed & 22 failed.
We (+@dumbbell) removed the tests which fail more often than they pass.
If anyone wants to add them back, please rewrite them as in the current
form they simply don't work. Alternatively, they are hinting to a
failure in the system which needs addressing. Both are out of the
current scope.
cc @michaelklishin
Signed-off-by: Gerhard Lazu <gerhard@lazu.co.uk>
|
| | |
| |
| |
| |
| |
| |
| |
| |
| |
| |
| |
| |
| |
| |
| |
| |
| |
| |
| |
| |
| |
| |
| |
| |
| |
| |
| |
| |
| |
| |
| |
| |
| |
| |
| |
| |
| |
| |
| |
| | |
Conclusion of the following conversation:
- Gerhard Lazu: what should we do about these recurring test failures?
https://github.com/rabbitmq/rabbitmq-server/runs/548636403?check_suite_focus=true#step:6:185
== publisher_confirms_parallel_SUITE ==
publisher_confirms_parallel_SUITE > publisher_confirm_tests > mirrored_queue > publisher_confirms
#1. {error,{noproc,{gen_server,call,
[<0.393.0>,
{wait_for_confirms,5000},
infinity]}}}
publisher_confirms_parallel_SUITE > publisher_confirm_tests > mirrored_queue > confirm_nack
#1. {error,
{{badmatch,received_ack_instead_of_nack},
[{publisher_confirms_parallel_SUITE,confirm_nack,1,
[{file,"test/publisher_confirms_parallel_SUITE.erl"},
{line,247}]},
{test_server,ts_tc,3,[{file,"test_server.erl"},{line,1562}]},
{test_server,run_test_case_eval1,6,
[{file,"test_server.erl"},{line,1080}]},
{test_server,run_test_case_eval,9,
[{file,"test_server.erl"},{line,1012}]}]}}
- Karl Nilsson: I always suspected that test to be racy
- Gerhard Lazu: In which case, I would like to delete it. Does someone want to fix it instead?
- Karl Nilsson: How do we know the queue exits after the channel processing the publish but before the queue mirrors sending all acks the and ch returning with ack
- Michael Klishin: If there is no way to make it not racy or something like Jepsen would make more sense to test this kind of thing, we can simply delete the test
- Karl Nilsson: Great point Michael, the Jensen suite should provide some coverage of this.
- Karl Nilsson: I think we are safe to remove these tests, at least the ones using mirroring
- Michael Klishin: Please go ahead and remove them
Signed-off-by: Gerhard Lazu <gerhard@lazu.co.uk>
|
| | |
| |
| |
| |
| |
| |
| |
| |
| |
| |
| |
| |
| |
| |
| |
| | |
https://github.com/rabbitmq/rabbitmq-server/runs/522909848?check_suite_focus=true#step:4:512
simple_ha_SUITE > cluster_size_3 > overflow_reject_publish > rejects_survive_stop
#1. {error,{error,received_both_acks_and_nacks}}
- It's not helping anyone in the current state
- I don't have enough context to be able to fix it
- I need to stay focused on the current task, cannot afford to context switch
- Feel free to fix it if it's important, otherwise leave it deleted
cc @michaelklishin @dumbbell
Signed-off-by: Gerhard Lazu <gerhard@lazu.co.uk>
|
| | |
| |
| |
| |
| |
| |
| |
| |
| |
| |
| |
| |
| |
| |
| |
| |
| |
| |
| |
| |
| |
| |
| |
| |
| |
| |
| |
| |
| |
| |
| |
| |
| |
| | |
https://github.com/rabbitmq/rabbitmq-server/runs/522627483?check_suite_focus=true#step:4:769
backing_queue_SUITE > backing_queue_tests > backing_queue_embed_limit_0 > variable_queue_default > variable_queue_fold
#1. {error,
{error,
{case_clause,undefined},
[{file_handle_cache,'-partition_handles/1-fun-0-',2,
[{file,"src/file_handle_cache.erl"},{line,770}]},
{file_handle_cache,get_or_reopen,1,
[{file,"src/file_handle_cache.erl"},{line,709}]},
{file_handle_cache_stats,timer_tc,1,
[{file,"src/file_handle_cache_stats.erl"},{line,63}]},
{file_handle_cache_stats,update,2,
[{file,"src/file_handle_cache_stats.erl"},{line,49}]},
{file_handle_cache,with_handles,3,
[{file,"src/file_handle_cache.erl"},{line,666}]},
{rabbit_msg_store,read_from_disk,2,
[{file,"src/rabbit_msg_store.erl"},{line,1251}]},
{rabbit_msg_store,client_read3,3,
[{file,"src/rabbit_msg_store.erl"},{line,674}]},
{rabbit_msg_store,safe_ets_update_counter,5,
[{file,"src/rabbit_msg_store.erl"},{line,1306}]}]}}
- It's not helping anyone in the current state
- I don't have enough context to be able to fix it
- I need to stay focused on the current task, cannot afford to context switch
- Feel free to fix it if it's important, otherwise leave it deleted
cc @michaelklishin @dumbbell
Signed-off-by: Gerhard Lazu <gerhard@lazu.co.uk>
|
| | |
| |
| |
| |
| |
| |
| |
| |
| |
| |
| |
| |
| |
| |
| |
| |
| |
| |
| |
| |
| | |
Gets aborted after 30 minutes:
https://github.com/rabbitmq/rabbitmq-server/runs/522332298?check_suite_focus=true#step:4:479
queue_master_location_SUITE > cluster_size_3 > calculate_client_local
#1. {skip,
{failed,
{queue_master_location_SUITE,init_per_testcase,
{timetrap_timeout,1800000}}}}
- It's not helping anyone in the current state
- I don't have enough context to be able to fix it
- I need to stay focused on the current task, cannot afford to context switch
- Feel free to fix it if it's important, otherwise leave it deleted
cc @michaelklishin @dumbbell
Signed-off-by: Gerhard Lazu <gerhard@lazu.co.uk>
|
| | |
| |
| |
| |
| |
| |
| |
| |
| |
| |
| |
| |
| |
| |
| |
| |
| |
| |
| |
| |
| |
| |
| |
| | |
https://github.com/rabbitmq/rabbitmq-server/runs/521986549#step:4:468
confirms_rejects_SUITE > parallel_tests > mixed_dead_alive_queues_reject
#1. {error,
{expecting_nack_got_ack,
[{confirms_rejects_SUITE,mixed_dead_alive_queues_reject,1,
[{file,"test/confirms_rejects_SUITE.erl"},{line,158}]},
{test_server,ts_tc,3,[{file,"test_server.erl"},{line,1562}]},
{test_server,run_test_case_eval1,6,
[{file,"test_server.erl"},{line,1080}]},
{test_server,run_test_case_eval,9,
[{file,"test_server.erl"},{line,1012}]}]}}
- It's not helping anyone in the current state
- I don't have enough context to be able to fix it
- I need to stay focused on the current task, cannot afford to context switch
- Feel free to fix it if it's important, otherwise leave it deleted
cc @michaelklishin @dumbbell
Signed-off-by: Gerhard Lazu <gerhard@lazu.co.uk>
|
| | |
| |
| |
| |
| |
| |
| |
| |
| |
| |
| |
| |
| |
| |
| |
| | |
https://github.com/rabbitmq/rabbitmq-server/runs/521844674?check_suite_focus=true#step:4:468
gm_SUITE > member_death
#1. {error,{thrown,timeout_waiting_for_death_3_2}}
- It's not helping anyone in the current state
- I don't have enough context to be able to fix it
- I need to stay focused on the current task, cannot afford to context switch
- Feel free to fix it if it's important, otherwise leave it deleted
cc @michaelklishin @dumbbell
Signed-off-by: Gerhard Lazu <gerhard@lazu.co.uk>
|
| | |
| |
| |
| |
| |
| |
| |
| |
| |
| |
| | |
- It's not helping anyone in the current state
- I don't have enough context to be able to fix it
- I need to stay focused on the current task, cannot afford to context switch
- Feel free to fix it if it's important, otherwise leave it deleted
cc @michaelklishin @dumbbell
Signed-off-by: Gerhard Lazu <gerhard@lazu.co.uk>
|
| |\ \
| | |
| | | |
Add docs for check_if_node_is_mirror_sync_critical and check_if_node_…
|
| |/ /
| |
| |
| |
| |
| | |
check_if_node_is_quorum_critical
Part of rabbitmq/rabbitmq-cli#389
|
| |\ \
| | |
| | | |
Wait until node detects new cluster configuration
|
| |/ /
| |
| |
| |
| |
| |
| | |
CI has failed with an mnesia error where the rabbit_queue table
doesn't exist. The actual logs don't show any error on the remaining
node so let's assume that is mnesia detecting the other node going
down. This really shouldn't happen, but I can't reproduce it either.
|
| |/ |
|
| |\
| |
| | |
unit_log_management_SUITE: Simplify code of `log_file_fails_to_initialise_during_startup`
|
| | |
| |
| |
| |
| |
| |
| | |
`log_file_fails_to_initialise_during_startup`
Also, add more log messages to help us debug this testcase when it
fails.
|
| |/
|
|
|
|
|
|
|
|
|
|
|
|
| |
Some of them we run up to 20 times (!!!) to make sure that they succeed.
- They are not helping anyone in the current state
- I don't have enough context to be able to fix them
- I need to stay focused on the current task, cannot afford to context switch
- Feel free to fix it if it's important, otherwise leave it deleted
cc @michaelklishin @dumbbell
Signed-off-by: Gerhard Lazu <gerhard@lazu.co.uk>
(cherry picked from commit a835d3271680ad6db5663f504f08fd0db4ee21c2)
|
| |\
| |
| | |
rabbit_feature_flags: Multiple fixes and optimizations to get rid of race conditions
|
| | |
| |
| |
| |
| | |
We need to handle concurrent calls to this function to avoid any issues
with parallel read/modify/write operations.
|
| | |
| |
| |
| |
| |
| |
| |
| |
| |
| |
| |
| | |
Before we query all the details needed to generate a new registry, we
save the version of the currently loaded one, using the `vsn` Erlang
module attribute added by the compiler.
When the new registry is ready to be loaded, we verify again the version
of the loaded one: if it differs, it means a concurrent process reloaded
the registry. In this case, we restart the entire regen procedure,
including the query of fresh details. The goal is to avoid any loss of
information from the existing registry.
|
| | |
| |
| |
| |
| |
| |
| |
| |
| |
| |
| |
| |
| | |
By calling code:delete() ourselves, we created a time window where there
is no registry module loaded between that call and the following
code:load_binary().
This meant that a concurrent access to the registry would trigger a load
of the initial uninitialized registry module from disk. That module
would then trigger a reload itself, leading to possible deadlock.
In fact, in code:load_binary(), the code server already takes care of
replacing the module atomically. We don't need to do anything.
|
| | | |
|
| | |
| |
| |
| |
| | |
Now that the feature flags subsystem uses its own log category, we need
to configure the corresponding handler.
|
| | |
| |
| |
| |
| |
| |
| |
| |
| |
| | |
As of this commit, it show two issues with the current implementation of
the feature flags registry reloading:
* There is a small time window where there is no registry module loaded
while it is replaced.
* The newer registry might be initialized with older data.
Follow-up commits will address them.
|
| | |
| |
| |
| |
| |
| | |
Why did I write this in French in the first place?
While here, fix a typo in the comment nearby.
|
| |/
|
|
|
|
|
| |
... from unknown applications were discovered.
In particular, this was triggerring more queries to remote nodes: more
time spent to regen the same registry and potential for errors.
|
| |\
| |
| | |
Fixes race conditions on rabbit_fifo_int
|
| | |
| |
| |
| |
| |
| |
| |
| | |
A few testcases are time dependant. Instead of waiting a
predetermined amount of time for the ra events, this PR waits
for a specific number of events. This should remove most false
negatives detected on CI, even though we still have timeouts -
but much longer!
|
| |\ \
| |/
|/| |
Fix flaky test - connection tracking is asynchronous
|
| | | |
|
| |/ |
|
| |
|
|
|
|
|
|
|
| |
Such cases are best tested using other tools. These cases
are highly flaky and prevent pipeline progress. We believe
it's the tests that are timing-sensitive.
These may be replaced with e.g. additional Jepsen tests as needed
in the future.
|
| |\
| |
| | |
Remove Ra segment_max_entries override
|
| | |
| |
| |
| |
| | |
So that it uses Ra's internal default of 4096 instead which is safer for
larger message sizes.
|
| |\ \
| | |
| | |
| | |
| | | |
rabbitmq/handle-deadlocks-in-peer_discovery_classic_config_SUITE
peer_discovery_classic_config_SUITE: Handle dead-locks
|
| | | | |
|