idea is a simple OpenAI-compatible API:
Access to open-weight models
zero retention of prompts and completions
no training on data
only retain metadata required for billing and operations: request ID, model, input/output token counts, latency, timestamp, etc.
no request/response in logs
Just here to check interest. I am not 100% sure if I can get compute to do it or not (may be i can just try with my 5090 first with a small model), but I just want to check and talk with people rather than just thinking in my head.