Threads, Timers and Work Queues in Zephyr


Zephyr RTOS, part 5. Doing several things at once on a single CPU: what a thread really is, how to create one, how big to make its stack, and the three kernel objects that let an interrupt hand work to a thread.
Parts 1 to 4 ended with a board that boots and a console that answers. Everything so far ran inside main(). This part is about doing more than one thing at a time, which is the whole reason to put an RTOS on a microcontroller in the first place.
The problem with one big loop
If you come from Arduino, this shape is familiar:
while (1) {
read_sensor();
update_control();
send_to_cloud();
}It works until one of those calls takes its time. Say the sensor wants reading every 10 ms, the control loop every 1 ms, and the cloud upload blocks for 300 milliseconds waiting for the network. For those 300 ms the control loop simply does not run. Your motor jerks, your sensor misses samples, and no amount of rewriting send_to_cloud() fixes the structure.
An RTOS splits that loop into pieces that can wait independently. The piece blocked on the network stops being the problem, because while it waits, the kernel runs something else.
What a thread actually is
A thread is an independent sequence of execution. Each one behaves as if it owns the CPU, and the kernel's scheduler is the referee deciding who actually gets it at each instant.
Two things make a thread:
A stack. Its own slice of RAM, for local variables and the chain of function calls it is inside. Every thread needs one, and sizing it is the fiddly part — see below.
An entry function. The C function it starts in, which always takes three
void *parameters whether you use them or not.
main() is itself a thread, created for you. That has a consequence worth stating plainly: returning from main() ends that one thread and nothing else. Your other threads keep running, which is why the shell stayed alive in part 4 even though main() returned immediately.
Creating a thread
The static form declares the thread, its stack and its start-up at compile time:
/* 1. How much stack, and at what priority */
#define SENSOR_STACK_SIZE 1024
#define SENSOR_PRIORITY 5
/* 2. The entry function: three void* parameters, used or not */
static void sensor_thread(void *p1, void *p2, void *p3)
{
ARG_UNUSED(p1); ARG_UNUSED(p2); ARG_UNUSED(p3);
while (1) {
LOG_INF("sensor: sampling");
k_msleep(100); /* gives the CPU up for 100 ms */
}
}
/* 3. name, stack size, entry, p1, p2, p3, priority, options, start delay */
K_THREAD_DEFINE(sensor_tid, SENSOR_STACK_SIZE, sensor_thread, NULL, NULL, NULL,
SENSOR_PRIORITY, 0, 0);That is the whole thing. sensor_tid becomes an identifier you can pass to other calls. The last argument is a start delay: 0 starts the thread at boot, and SYS_FOREVER_MS leaves it dormant until you call k_thread_start(sensor_tid).
There is a runtime form too, for threads whose existence you only decide while running:
static K_THREAD_STACK_DEFINE(sensor_stack, SENSOR_STACK_SIZE);
static struct k_thread sensor_thread_data;
k_tid_t tid = k_thread_create(&sensor_thread_data, sensor_stack,
K_THREAD_STACK_SIZEOF(sensor_stack),
sensor_thread, NULL, NULL, NULL,
K_PRIO_PREEMPT(5), 0, K_NO_WAIT);Two traps live in those two snippets. The stack must come from K_THREAD_STACK_DEFINE, never a plain uint8_t array, because the macro applies the alignment and guard rules the architecture needs. And the start delay is milliseconds in K_THREAD_DEFINE but a k_timeout_t in k_thread_create — the one place these two APIs disagree.
Priorities: the lower number wins
Zephyr does not treat threads equally, and the numbering runs against intuition: a lower number means a higher priority. Priority 2 beats priority 7.
The sign splits threads into two kinds:
Priority | Kind | Behaviour |
|---|---|---|
Negative, | Cooperative | Keeps the CPU until it blocks, yields or ends. Nothing in software can take it away. |
Zero or positive, | Preemptible | Loses the CPU the instant a higher-priority thread becomes ready. |
Most application work is preemptible. Cooperative priorities are for short sequences that must not be interrupted by other threads — and they are a loaded gun, because a cooperative thread that never blocks never gives the CPU back.
Which brings us to the rule that catches everyone once:
A thread only gives up the CPU when it blocks.
k_msleep,k_sem_take,k_msgq_get, a contendedk_mutex_lock— those block. A busywhile (1) { }does not, and it will starve every lower-priority thread on the system.
The four states a thread can be in
State | Meaning |
|---|---|
Running | It has the CPU right now |
Ready | It wants to run, but something higher-priority has the CPU |
Waiting | It is blocked on a timeout, a semaphore, a queue or another object |
Suspended | Something called |
And the calls that move threads between them, all usable from another thread:
k_thread_suspend(sensor_tid);
k_thread_resume(sensor_tid);
k_thread_abort(sensor_tid); /* last resort */
k_thread_join(&sensor_thread_data, K_SECONDS(1)); /* wait for it to finish */How big should the stack be?
This is the hardest number in the file, and guessing badly hurts in both directions. Too small and the thread writes past its stack into memory that is not its own — the symptom is a fault somewhere unrelated, or silent corruption. Too large and you waste RAM, the scarcest thing on a microcontroller.
Do not guess twice. Measure, with the thread analyzer:
CONFIG_THREAD_ANALYZER=y
CONFIG_THREAD_ANALYZER_AUTO=y
CONFIG_THREAD_ANALYZER_AUTO_INTERVAL=10
That one symbol pulls in everything it needs (INIT_STACKS, THREAD_MONITOR, THREAD_STACK_INFO), and AUTO starts a thread that prints a report every interval — ten seconds here — with one line per thread: its name, how many bytes of stack it has ever used, the size you gave it, and the percentage. Let the board run through its busiest path, then read the numbers. Under about half, shrink it; over about 80 %, grow it now.
If your SoC supports it, CONFIG_HW_STACK_PROTECTION=y turns an overflow into a clean fault instead of a mystery. On RISC-V that needs PMP support in the SoC, so check before relying on it.
From the shell, kernel thread list and kernel thread stacks answer the same question interactively.
Not everything needs a thread
A thread is not free: its stack is RAM reserved whether it is busy or asleep, and the scheduler has one more candidate to weigh. For a job that takes microseconds every half second, three lighter objects do better — a timer, a work item and a message queue.
k_timer: do something every N milliseconds
static void on_tick(struct k_timer *timer)
{
/* interrupt context */
}
K_TIMER_DEFINE(tick_timer, on_tick, NULL); /* third arg: optional stop function */
k_timer_start(&tick_timer, K_MSEC(500), K_MSEC(500)); /* first at 500 ms, then every 500 ms */
k_timer_start(&tick_timer, K_MSEC(500), K_NO_WAIT); /* one-shot */
k_timer_stop(&tick_timer);What that looks like over time, for the periodic form:
t = 0 ms k_timer_start() registers a deadline and returns immediately
nothing of yours runs, and no CPU is spent waiting
t = 500 ms the timer interrupt fires -> on_tick() runs
t = 1000 ms on_tick() runs
t = 1500 ms on_tick() runs ... until k_timer_stop()Nothing polls during those gaps. The kernel programs the hardware timer for the next deadline, so between expiries the CPU is free for other threads, or asleep in the idle thread.
A periodic timer measures each period from the previous expiry, so periods do not drift even if one expiry is serviced late. The catch is the context: the expiry function runs in an interrupt, so it does one small thing — set an atomic flag, give a semaphore, submit a work item, put a message on a queue — and returns. It may not block or take a mutex, because an interrupt is not a thread: there would be nothing for the kernel to put to sleep.
k_work: hand the job to a thread
A work item is a function plus the bookkeeping to queue it. Submitting is safe from an interrupt, and the handler runs in the system workqueue's thread, where blocking and logging are fine again:
static void report(struct k_work *work)
{
/* thread context */
}
K_WORK_DEFINE(report_work, report);
k_work_submit(&report_work); /* safe from an ISR */
The handover, when the timer expiry above submits it:
t = 500 ms interrupt: k_work_submit() puts the item in the queue and returns
the ISR ends here — your handler has not run yet
t = 500 ms the workqueue thread is ready again, the scheduler runs it
t = 500 ms report() runs, in thread context, free to log and to blockThose three steps happen back to back when the queue is idle. If the workqueue is busy with another handler, yours simply waits its turn — the delay is however long the handler in front takes.
Two behaviours to know. The workqueue runs handlers one at a time, in order, so a handler that blocks holds up everyone else's work — including other subsystems, since the system queue is shared. And submitting an item that is already queued does nothing: it runs once. If every event must be counted, count it where you submit, not in the handler.
For "do it after a delay" there is a delayable variant, with two ways to start it:
K_WORK_DELAYABLE_DEFINE(debounce_work, handler);
k_work_schedule(&debounce_work, K_MSEC(20)); /* first call wins: keeps an existing deadline */
k_work_reschedule(&debounce_work, K_MSEC(20)); /* last call wins: always moves the deadline */The difference is what a second call does while a run is already pending. Take a bouncing contact that gives you edges at 0, 5 and 9 ms, with a 20 ms delay:
k_work_reschedule | k_work_schedule
---------------------------+---------------------------
t = 0 ms deadline 20 ms | t = 0 ms deadline 20 ms
t = 5 ms deadline 25 ms | t = 5 ms ignored, deadline stays 20 ms
t = 9 ms deadline 29 ms | t = 9 ms ignored, deadline stays 20 ms
t = 29 ms handler runs | t = 20 ms handler runsk_work_reschedule runs the handler 20 ms after the last edge, which is what debounce means: wait for the input to go quiet. k_work_schedule runs it 20 ms after the first, which is what you want when a burst of changes should produce exactly one flush. Either way the handler runs once, not three times.
If your handlers block or run long, give them a queue of their own with k_work_queue_start() instead: same API, your own thread and priority.
k_msgq: when data travels
A work item carries no payload. When the handover does, declare a message queue: a ring buffer of fixed-size messages, copied in and out.
K_MSGQ_DEFINE(reading_q, sizeof(struct reading), 8, 4); /* message size, count, alignment */
k_msgq_put(&reading_q, &r, K_NO_WAIT); /* producer, ISR-safe; fails when full */
k_msgq_get(&reading_q, &r, K_FOREVER); /* consumer, in its own thread */The two sides meet in the middle:
t = 0 ms consumer calls k_msgq_get(..., K_FOREVER)
queue is empty -> the thread leaves the ready queue and sleeps
t = 500 ms interrupt: k_msgq_put() copies the message in and returns
the kernel marks the consumer ready; the ISR ends
t = 500 ms the consumer wakes, k_msgq_get() returns 0, and it has the dataA message put while nobody is waiting is not lost: it stays in the ring buffer, and the next k_msgq_get() returns it at once. From an interrupt the timeout must be K_NO_WAIT, which means you must decide what happens when the queue is full — usually dropping the sample and counting the drop. k_msgq_num_used_get() tells you how close you are, which is how you find out a producer is outrunning a consumer before data starts disappearing.
The near relative is k_fifo, which passes pointers to buffers you own instead of copying them.
k_sem and k_mutex: signal, or protect
These two look alike — both can make a thread wait — but they answer different questions. k_sem answers "has something happened?": a counter with no owner, which anyone may give, including an interrupt. k_mutex answers "may I touch this data?": a lock with an owner, which only the locking thread may unlock.
|
| |
|---|---|---|
Really is | a counter | a lock with an owner |
From an interrupt |
| never: mutexes may not be locked in ISRs |
Priority inheritance | no | yes |
Reach for it when | signalling between contexts | protecting shared data |
Do not use a binary semaphore as a lock. With no owner the kernel cannot raise a holder's priority, so a medium-priority thread can preempt the holder while a high-priority thread waits on it — unbounded priority inversion. A mutex is built for exactly that case. And for a single word, a counter or a flag, atomic_t beats both: no waiting, and usable from an interrupt.
All three handovers side by side, with the rule that applies to each:

The whole thing in one file
A timer sampling in an interrupt, a queue carrying the samples out, a thread consuming them, and a work item reporting every tenth sample:
#include <zephyr/kernel.h>
#include <zephyr/logging/log.h>
LOG_MODULE_REGISTER(demo, LOG_LEVEL_INF);
struct reading {
uint32_t time_ms;
uint32_t value;
};
K_MSGQ_DEFINE(reading_q, sizeof(struct reading), 8, 4);
/* Thread context: logging and blocking are fine here. */
static void report(struct k_work *work)
{
ARG_UNUSED(work);
LOG_INF("report: %u readings waiting", k_msgq_num_used_get(&reading_q));
}
K_WORK_DEFINE(report_work, report);
/* Interrupt context: no blocking, K_NO_WAIT only, return fast. */
static void on_tick(struct k_timer *timer)
{
ARG_UNUSED(timer);
static uint32_t count;
struct reading r = { .time_ms = k_uptime_get_32(), .value = count++ };
(void)k_msgq_put(&reading_q, &r, K_NO_WAIT);
if ((count % 10U) == 0U) {
k_work_submit(&report_work);
}
}
K_TIMER_DEFINE(tick_timer, on_tick, NULL);
/* The consumer blocks here until a reading arrives: no polling, no CPU burnt. */
static void consumer(void *p1, void *p2, void *p3)
{
ARG_UNUSED(p1); ARG_UNUSED(p2); ARG_UNUSED(p3);
struct reading r;
while (k_msgq_get(&reading_q, &r, K_FOREVER) == 0) {
LOG_INF("reading %u at %u ms", r.value, r.time_ms);
}
}
K_THREAD_DEFINE(consumer_tid, 1024, consumer, NULL, NULL, NULL, 5, 0, 0);
int main(void)
{
LOG_INF("demo starting");
k_timer_start(&tick_timer, K_MSEC(500), K_MSEC(500));
return 0;
}Drop it into the app_template project from part 4 with CONFIG_LOG=y in prj.conf, and build and flash it the same way. Mine builds to a 136 kB image for esp32p4_wifi6/esp32p4/hpcore. On the console you get a reading every half second, and after every tenth one, a report line from the workqueue.
Three contexts are at work in those sixty lines: an interrupt that only hands things over, a workqueue thread that reports, and a thread of your own that sleeps until there is something to do. That is the shape of nearly every Zephyr application.
Which one to reach for
You want | Use |
|---|---|
Run something in a thread, from an interrupt |
|
…after a delay, or debounced |
|
A periodic tick |
|
Hand over data |
|
Signal "something happened" |
|
Protect one word |
|
Protect a multi-step update |
|
Long or blocking work, or its own priority | a thread, or your own workqueue |
What's next
Part 6 puts these to work on something real: turning button handling into an out-of-tree Zephyr module, with its own Kconfig options, a state machine that can be tested without hardware, and buttons read from the devicetree.
Resources
Threads: docs.zephyrproject.org/latest/kernel/services/threads
Timers: docs.zephyrproject.org/latest/kernel/services/timing/timers.html
Workqueue threads: docs.zephyrproject.org/latest/kernel/services/threads/workqueue.html
Message queues: docs.zephyrproject.org/latest/kernel/services/data_passing/message_queues.html
Thread analyzer: docs.zephyrproject.org/latest/services/debugging/thread-analyzer.html
Found a mistake or something unclear? Let me know in the comments below, and I will correct the post.