RL Pipeline¶
Virne models NFV-RA solution construction as a sequential decision problem and provides reusable components for training and evaluating reinforcement learning (RL) solvers.
Virne’s reusable RL pipeline connects environment interaction, rollout collection, policy optimization, checkpointing, and evaluation.¶
NFV-RA as an MDP¶
For the common node-by-node formulation, the Markov Decision Process (MDP) is defined by \((S, A, P, R, \gamma)\):
State \(s_t\): the PN resources, the current VN request, and the partial mapping at decision step \(t\).
Action \(a_t\): usually the physical node selected for the current virtual node. Invalid actions can be masked.
Transition \(P\): applies placement and routing checks, updates the partial solution, and advances to the next virtual node or a terminal state.
Reward \(R(s_t, a_t)\): provides intermediate and/or terminal feedback based on feasibility and resource efficiency.
Discount \(\gamma\): balances immediate and future rewards.
The policy \(\pi_\theta(a_t \mid s_t)\) is trained to maximize expected discounted return:
How the Pipeline Maps to Code¶
BaseSystemgenerates VN arrival events and passes each instance to the selected solver.An instance-level environment constructs a solution one action at a time and delegates feasibility checks to the shared controller.
The feature constructor converts the current graph state into tensors; the policy network selects an action, optionally using an action mask.
RolloutBufferstores actions, rewards, values, and terminal flags.RLSolverupdates the policy, saves checkpoints, switches to evaluation mode, and solves the configured simulation.
This design keeps simulation, feasibility checking, policy architecture, and training logic separable. Different RL solvers can therefore share the same network scenarios and evaluation pipeline.
Instance-level RL environments follow the Gymnasium 1.3 API: reset()
returns (observation, info) and step() returns (observation, reward,
terminated, truncated, info). Virne’s solution-building episodes currently
end through natural termination, so they return truncated=False. Training
code still keeps both flags distinct so future external time limits can
bootstrap value estimates correctly.
Run a Minimal Training Check¶
The following CPU command trains ppo_dual_gat+ for one epoch on ten VN
requests, saves a checkpoint, and then evaluates the trained policy:
virne \
solver.solver_name=ppo_dual_gat+ \
v_sim_setting.num_v_nets=10 \
training.num_train_epochs=1 \
training.use_cuda=false \
rl.target_steps=32 \
'logger.backends=[console]'
Note
This is a pipeline smoke test, not a meaningful benchmark. Research results require larger training and evaluation sets, controlled seeds, and the protocol described in the benchmark paper.
Discrete Off-policy Solvers¶
Virne also registers dqn_mlp+, double_dqn_mlp+, and ddpg_mlp+ for
the physical-node action space. They use persistent replay memory and target
networks; invalid physical-node actions are masked during both exploration and
Bellman target calculation. Their settings live under rl.dqn and
rl.ddpg in virne/configs/learning.yaml.
ddpg_mlp+ is a discrete-action adaptation, not vanilla continuous-action
DDPG. Its actor emits masked node logits, its critic estimates one Q-value per
physical node, and the actor is optimized through a differentiable categorical
relaxation. This distinction should be retained when reporting the algorithm
in experiments.
The final checkpoint is written to:
results/virne/ppo_dual_gat+/<run-id>/models/model.pkl
Key Training Settings¶
Setting |
Purpose |
|---|---|
|
Selects the registered RL solver and policy architecture. |
|
Trains before evaluation when greater than zero. |
|
Select CPU or a CUDA device. |
|
Controls intermediate checkpoint frequency. |
|
Sets the rollout size used to trigger an update. |
|
Control discounted returns and generalized advantage estimation. |
|
Selects the reward calculation strategy and intermediate reward. |
|
Selects the state features supplied to the policy. |
|
Enables or disables invalid-action masking. |
Defaults and the remaining optimizer and network settings live in
virne/configs/main.yaml and virne/configs/learning.yaml.
Evaluate a Saved Model¶
Set training epochs to zero and pass an absolute checkpoint path:
virne \
solver.solver_name=ppo_dual_gat+ \
training.num_train_epochs=0 \
solver.pretrained_model_path=/absolute/path/to/model.pkl
Without solver.pretrained_model_path, setting
training.num_train_epochs=0 evaluates a randomly initialized RL policy.
See Learning-based Solver for the RL solver families implemented in Virne and Solver Registry for the complete registry of valid command names.