Skip to content
MARL in Cooperative Environments
Edit this page

Wireless Allocation Experiments

10 min read

This Colab activity integrates coordination, communication, and partner adaptation in the supplied wireless environment. Short experiments move from manual allocation and independent learning to coordinated policies, what communication costs, a traffic shift, and an unfamiliar access point. You will finish by comparing every system in one table and defending a design choice. The environment and baselines come from the Python package, so the notebook holds the experiment and not the infrastructure.

wireless_network_resource_allocation_marl_lab.ipynb
Time
60 to 90 minutes
Compute
Free Colab CPU is enough; no GPU

Before any learning, you choose joint actions yourself and watch the network respond. Then the notebook enumerates all 81 allocations to show whether you found the best one.

Four access points on three channels means one pair must share, so the question is never whether to share. It is which pair.

Two quantities decide it, and they do not carry equal weight:

  • Demand. Useful throughput is min⁡(di,ratei)\min(d_i, \text{rate}_i). A light access point does not notice a slower channel, because its throughput was capped by its own demand anyway.
  • Distance. Coupling falls off as 1/(1+(d/d0)2)1/(1 + (d/d_0)^2). In the supplied layout AP0 and AP1 couple at 0.671 while AP0 and AP3 couple at 0.059, an order of magnitude apart.
AllocationRewardWhat is sharing
[0, 0, 0, 0]3.51everybody on one channel
[2, 2, 1, 0]5.73a heavy access point with its closest neighbour
[0, 1, 1, 2]7.92the two light access points
Best of all 817.92the same allocation

One episode, seed 42, where AP0 and AP3 are the busy pair. It shows the shape of the problem rather than an average.

Demand decides which pair shares. Distance decides what it costs when a busy access point has to. Pairing the two light ones is optimal even when they are the closest pair in the network: a channel shared at coupling 0.671 still delivers 1.20, which is more than the 0.8 they wanted.

Two rules that do not learn, then two that do. The greedy rule picks the channel with the best quality minus the interference it measured, which is individually sensible.

SystemRewardThroughputAvoidable interferenceCollision rate
Random6.006.190.880.47
Greedy local3.373.962.831.00
Independent7.177.240.220.14
Coordinated (VDN)7.567.620.150.07
Ceiling7.90

Two results here, and the first one surprises almost everybody.

Greedy local scores below random. Every decision it makes is defensible in isolation and it finishes at 3.37 against random’s 6.00. Its collision rate is 1.00: every access point is in harmful interference on every step. Because all four run the identical deterministic rule on a similar observation, they move to the same channel together. Random never synchronises, and that alone is worth 2.6 reward here.

Coordination is worth 0.39. Independent learning reaches 7.17, value decomposition 7.56, and the ceiling is 7.90. Execution is identical between the two; only the training signal changes. Avoidable interference falls from 0.22 to 0.15 and the collision rate halves.

5 · What communication buys, and what it costs

Section titled “5 · What communication buys, and what it costs”

Give each access point one bit carrying its demand level, so the team sends four messages per step. That mapping is a design decision, not a property of the problem.

The interesting question is not whether the bits are useful. It is whether they are useful enough.

SystemRewardThroughputMessages
Coordinated7.567.620
Coordinated + comm, free7.647.724
Coordinated + comm, price 0.057.437.704

Talking raises throughput, from 7.62 to 7.70. And it lowers the team reward, from 7.56 to 7.43, because four messages at 0.05 each is a bill of 0.20 per step against a gain of about 0.08.

Sweep the price and find where talking stops paying.

Team reward against Price per message. 0.00: 7.6, 0.02: 7.5, 0.05: 7.4, 0.10: 7.2, 0.25: 6.7 01.93.85.77.60.000.020.050.100.25Price per messageTeam reward

The silent coordinated system scores 7.56, so that is the line communication has to beat.

Price per messageRewardAgainst silence (7.56)
0.007.64talking wins, barely
0.027.53silence wins
0.057.43silence wins
0.107.15silence wins
0.256.71silence wins

You can predict the crossover before running the sweep. Communication bought about 0.08 reward for four messages, so it breaks even near λ≈0.08/4=0.02\lambda \approx 0.08/4 = 0.02, and the sweep puts it there.

Being able to compute where a protocol stops paying is the practical skill this section is for. Note also how small the free-communication gain is: 0.08 on a reward of 7.6, which is the same order as the spread across training seeds. A single seed could have shown this protocol helping or hurting.

Train under skewed demand, where two access points are busy, then evaluate under hotspot, where one saturates and the rest go quiet. No retraining.

SystemTrained regimeHotspotDropGap to hotspot ceiling
Independent7.176.600.570.37
Coordinated7.566.870.690.10
Coordinated + comm7.433.623.813.36
Ceiling7.906.98

Read the last column, not the drop column. The hotspot regime is simply harder: its ceiling is 6.98 against 7.90. The coordinated system’s 0.69 drop leaves it 0.10 from optimal in the new regime, so almost all of that drop is the regime changing rather than the policy failing.

The communicating system is a different story. It falls to 3.62, below random.

That single result carries all three chapters at once. Conditioning on more information is a coordination decision, a communication decision, and an adaptation risk, and here it is the same herd failure that sank the greedy rule in section 4.

Replace one access point with equipment from another operator: it takes the best local channel regardless of anyone else. You did not train it.

SystemFamiliarWith a strangerDrop
Independent7.176.680.49
Coordinated7.567.290.27
Coordinated + comm7.436.550.88

Here the coordinated system is both the best and the most robust, and the talking system degrades most. That is the opposite ordering from section 7, where the smallest drop belonged to the weakest system.

The talking system learned to depend on bits arriving. Drop some, with no retraining.

Probability a message is lostRewardAgainst silence (7.56)
0.07.43silence already wins
0.16.85
0.35.99
0.55.19below random (6.00)

At half the messages lost the system scores below a random allocation. A lost message reads as a quiet neighbour here, so loss does not merely remove information: it supplies confident wrong information.

The notebook finishes by assembling everything into the table you submit.

SystemRewardThroughputCollision rateMessagesHotspotWith stranger
Random6.006.190.470
Greedy local3.373.961.000
Independent7.177.240.1406.606.68
Coordinated7.567.620.0706.877.29
Coordinated + comm7.437.700.1843.626.55
Ceiling7.906.98

The only substantial writing the lab asks for.

Which system would you deploy if communication bandwidth is limited and neighbouring access points may change?

Three to four sentences. Name the trade-off you are accepting, and point at what in the table supports it.

There is no single correct answer, but the tables do contain a tension worth resolving out loud. The talking system has the highest throughput and the lowest reward, needs bandwidth the premise says is scarce, and is the one that collapses when traffic moves. The silent coordinated system is the best on every column that charges for something. Saying why you would still consider the talking system, or why you would not, is the answer.

Knowledge check

Communication raised throughput from 7.62 to 7.70 and lowered team reward from 7.56 to 7.43. What does that pair of numbers establish?

Select one answer.

Run the experiments:Open in Colab

Not recall. Transfer.

Coordinate. You reasoned from local channel decisions to a joint network outcome, found the optimum by hand, measured a deterministic local rule scoring below random, and separated what independent learning achieves from what value decomposition adds on top of it.

Communicate. You measured what the information was worth, priced it per message, computed the break-even price before running the sweep, and found a protocol that improves the task measure while lowering the objective.

Adapt. You measured degradation under changed traffic and under an access point you did not train, read both against the new regime’s ceiling rather than against the drop, found the two tables disagree about which system is robust, and diagnosed a collapse that individual state coverage could not explain.

Next: the Final Project, where nobody supplies the environment either.