1
Fork 0

fix: update file link

This commit is contained in:
Tibo De Peuter 2026-04-08 15:43:12 +02:00
parent 16cc86b30c
commit 67027e8905
Signed by: tdpeuter
SSH key fingerprint: SHA256:u/h/LVoqKF1Iz02uOyxe6hcjmoZASCGV2HM0TG9ZMoU
2 changed files with 1 additions and 4 deletions

View file

@ -2,11 +2,11 @@
## Design & architecture ([`/design`](./design/))
- [Actor/critic architecture](./design/actor-critic.md): Description of the actor-critic pipeline.
- [Communication](./design/communication.md): Message propagation, Nerve-Net style.
- [Controllers](./design/controllers.md): Macroscopig brain toplogy, centralized, arm-level, segment-level.
- [Input/output](./design/input_action_spaces.md): Description of the model's input and output.
- [Learning algorithm](./design/learning_algorithm.md): RL techniques, i.e. PPO.
- [MLP architecture](./design/mlp_architecture.md): Module wiring, state/action spaces, and actor-critic pipeline.
- [Reward function](./design/learning_algorithm.md): Goals, fitness tracking, and reward structures.
## API reference ([`/api`](./api/))

View file

@ -20,7 +20,6 @@ This pipeline treats the agent as a single entity and uses standard Proximal Pol
Our policy and value networks use separate input networks/feature extractors as advised by the SEL3 course assistants and the blog. For continuous actions this should allow better learning at a small cost.
```mermaid
%%{ init: { 'flowchart': {'defaultRenderer': 'elk' } } }%%
graph TD
Obs([Global Observation])
@ -69,7 +68,6 @@ critic for all nodes at once, for the following reasons:
mathematically equivalent.
```mermaid
%%{ init: { 'flowchart': {'defaultRenderer': 'elk' } } }%%
graph TD
Obs([Local Observation])
@ -96,7 +94,6 @@ graph TD
Crit --> OutCrit
Prop -.->|"message passing"|Prop
```
## Implementation Details (Network Depth)