You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
Environment: queue master (07fd732), Tarantool 3.8.0 (also the same code in 1.5.0). memtx.
Related: #238 asks for an API to restart the TTL fiber. This issue is about what happens today when the fiber dies: it cannot be restarted, and for fifottl/limfifottl the queue state machine gets stuck in ENDING on the next RO switch, so the queue never comes back without an instance restart.
Summary
Any error inside a TTL iteration (a user on_task_change callback raising, an MVCC conflict, a vinyl read error, the nil dereference from the TTL branch, ...) makes the fiber return 1 (fifottl.lua#L164-L167) while self.fiber keeps pointing at the dead fiber. Consequences:
method.start() is a no-op because self.fiber is not nil (fifottl.lua#L415-L420). TTL/TTR/delay processing for the tube is dead.
On the next RO switch the state machine calls stop(), which does self.sync_chan:put(true) (fifottl.lua#L427). The channel is unbuffered and the only reader is dead (same root cause as fifottl/limfifottl: stop() blocks forever while RW, drop() leaks the TTL fiber #262), so the queue_state fiber blocks forever. The queue stays in ENDING, and put/take keep failing with queue is in ENDING state even after the instance is RW again.
For utubettl (buffered fiber.channel(1)) point 2 does not hang, but point 1 still holds until an RO->RW cycle.
Repro
localfiber=require('fiber')
box.cfg{}
localqueue=require('queue')
localfired=falselocaltube=queue.create_tube('t', 'fifottl', {
on_task_change=function(task, stat)
ifstat=='ttl' andnotfiredthenfired=trueerror('user callback failed once')
endend,
})
tube:put('a', {ttl=0.1})
fiber.sleep(0.3)
print('1. fiber after one callback error:', tube.raw.fiber:status())
localb=tube:put('b', {ttl=0.1})
fiber.sleep(0.3)
print('2. expired task still in space:', tube.raw.space:get(b[1]) ~=nil)
tube.raw:start()
print('3. fiber after start():', tube.raw.fiber:status())
box.cfg{read_only=true}
fiber.sleep(1)
print('4. queue.state() after RO switch:', queue.state())
box.cfg{read_only=false}
fiber.sleep(1)
print('5. queue.state() after RW switch:', queue.state())
print('6. put() after RW switch:', pcall(tube.put, tube, 'c', {ttl=0.1}))
Output on master:
1. fiber after one callback error: dead
2. expired task still in space: true
3. fiber after start(): dead
4. queue.state() after RO switch: ENDING
5. queue.state() after RW switch: ENDING
6. put() after RW switch: true nil -- put() returns nil, log: "put: queue is in ENDING state"
Expected
A failed iteration should not permanently kill TTL processing: either the fiber survives the error (log it and continue / restart with backoff), or at least self.fiber is cleared on exit so start() can create a new one, and stop() must never block on a channel nobody reads.
Suggested fix
Clear self.fiber in a guaranteed cleanup path when the fiber exits (checking it still points to the current fiber).
Make stop() tolerant of a dead fiber (check self.fiber:status() or use a buffered channel / fiber:cancel()).
Optionally restart the fiber after an error with a backoff instead of return 1; user callback errors could be pcall-ed in abstract.lua so a user bug cannot take the driver down.
Environment: queue master (07fd732), Tarantool 3.8.0 (also the same code in 1.5.0). memtx.
Related: #238 asks for an API to restart the TTL fiber. This issue is about what happens today when the fiber dies: it cannot be restarted, and for
fifottl/limfifottlthe queue state machine gets stuck inENDINGon the next RO switch, so the queue never comes back without an instance restart.Summary
Any error inside a TTL iteration (a user
on_task_changecallback raising, an MVCC conflict, a vinyl read error, the nil dereference from the TTL branch, ...) makes the fiberreturn 1(fifottl.lua#L164-L167) whileself.fiberkeeps pointing at the dead fiber. Consequences:method.start()is a no-op becauseself.fiberis not nil (fifottl.lua#L415-L420). TTL/TTR/delay processing for the tube is dead.stop(), which doesself.sync_chan:put(true)(fifottl.lua#L427). The channel is unbuffered and the only reader is dead (same root cause as fifottl/limfifottl: stop() blocks forever while RW, drop() leaks the TTL fiber #262), so thequeue_statefiber blocks forever. The queue stays inENDING, andput/takekeep failing withqueue is in ENDING stateeven after the instance is RW again.For
utubettl(bufferedfiber.channel(1)) point 2 does not hang, but point 1 still holds until an RO->RW cycle.Repro
Output on master:
Expected
A failed iteration should not permanently kill TTL processing: either the fiber survives the error (log it and continue / restart with backoff), or at least
self.fiberis cleared on exit sostart()can create a new one, andstop()must never block on a channel nobody reads.Suggested fix
self.fiberin a guaranteed cleanup path when the fiber exits (checking it still points to the current fiber).stop()tolerant of a dead fiber (checkself.fiber:status()or use a buffered channel /fiber:cancel()).return 1; user callback errors could be pcall-ed inabstract.luaso a user bug cannot take the driver down.