alvis/oO

Files

alvis faf44c18fc feat: ε-greedy v1 as active policy; dwell-time reward inference; offline sim framework

- Promote egreedy-v1 to active serving policy (ADR-0007): /score/egreedy + /reward/egreedy
  replaces linucb-v1 endpoints after offline sim shows +10.7% mean reward (−0.548 vs −0.606)
- Replace explicit helpful/not_helpful feedback with dwell-time inferred reward (inferReward):
  dismiss=−1.0, snooze=+0.1, done<15s=−0.3, done 15s–2min=+1.0, done 2–10min=+0.6, done>10min=+0.3
- Add ml/serving ε-greedy endpoints: /score/egreedy, /reward/egreedy, /stats/egreedy/{user_id}
  with d=7 feature vector (base 5 + sin/cos day-of-week encoding)
- Add offline simulation framework (ml/experiments/sim): rule/LLM/claude-code judges,
  two-phase score+reward, synthetic personas, task generator; results stored in sim_runs/sim_events
- Add /admin/simulations page: start runs, live-poll status, reward curve SVG, action/persona tables
- Fix egreedy day_of_week training skew: reward endpoint now uses actual dow instead of hardcoded 0
- Fix runner.py proxy bypass: httpx.Client(trust_env=False) for localhost ML calls
- Add dwellMs to TipFeedbackEvent contract and bus.test.ts fixture
- Schema: sim_runs, sim_events tables; tip_feedback gains dwell_ms, reward_milli columns
- ADR-0006: admin console framework; ADR-0007: egreedy-v1 policy selection rationale

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

2026-04-16 07:44:37 +00:00

api

feat: ε-greedy v1 as active policy; dwell-time reward inference; offline sim framework

2026-04-16 07:44:37 +00:00

auth

chore: scaffold oO monorepo with architecture, roadmap, and module stubs

2026-04-13 14:19:56 +00:00

events

chore: scaffold oO monorepo with architecture, roadmap, and module stubs

2026-04-13 14:19:56 +00:00

gateway

chore: scaffold oO monorepo with architecture, roadmap, and module stubs

2026-04-13 14:19:56 +00:00

integrations

refactor: architecture revision — modular monolith, auth-commit, event protobuf, privacy-from-day-0

2026-04-13 14:36:11 +00:00

notifier

chore: scaffold oO monorepo with architecture, roadmap, and module stubs

2026-04-13 14:19:56 +00:00

profile

chore: scaffold oO monorepo with architecture, roadmap, and module stubs

2026-04-13 14:19:56 +00:00

recommender

refactor: architecture revision — modular monolith, auth-commit, event protobuf, privacy-from-day-0

2026-04-13 14:36:11 +00:00

README.md

refactor: architecture revision — modular monolith, auth-commit, event protobuf, privacy-from-day-0

2026-04-13 14:36:11 +00:00

README.md

services/

Backend modules. Each owns a contract and ships its own README.md. In Phase 0 these are internal packages inside a single Node process (ADR-0003); they extract to their own processes as pressure justifies.

Dir	Role	Phase-0 shape	Extracts when
`gateway/`	BFF for clients; auth check; fan-out	in-proc router	never (stays as the edge)
`auth/`	Google OAuth (Apple in M1), sessions, JWT	Auth.js behind OIDC shape	mobile native ships (M3)
`profile/`	user profile, preferences, consents	in-proc module	team ownership diverges
`integrations/`	connectors + encrypted token vault	in-proc module	credential blast-radius isolation
`recommender/`	`POST /recommend` — policy-driven tip selection	in-proc; calls `ml/serving` from M1	scaling hotspot
`events/`	event bus + signal log	in-proc emitter (Phase 0); NATS (M1)	always a library + broker, not a service
`notifier/`	push/email delivery + quiet hours	in-proc; web push in M1	SLA divergence or mobile push scale

Contracts that cross module lines (HTTP or events) come from packages/shared-types/. In-module imports across modules are forbidden by import lint.