Very cool.
Curious if the results are better for certain domains vs others and if there's a sweet spot in terms of number of turns in the convos.
What have you seen deploying this in the wild so far?
So far, our early design partners have been running this on
(a) AI Based Ads creation and optimisation product [Multi Modal]
(b) Customer Support for eCommerce [Text Based] and
(c) our own AI trading system.
We’re actively looking for more use cases because we want to understand where the approach breaks down.
We haven't seen a clear domain specific sweet spot yet. What surprises us is that it becomes seems to get better as the complexity / turn count increases.
It adopts very well to multi turn sessions, scheduled runs and even trigger based Agentic systems that don't involve human in the loop.
We don’t yet have enough data to say something like “10–20 turns is optimal,” though. That’s one of the things we’re hoping to learn as we get it into more production systems.
Great question. It usually never happens. Our system is based on an adoption of Domino (systematic error/slice discovery), AgentBoard (trajectory/progress evaluation), τ-bench (goal/outcome correctness), MAST (failure taxonomies) and process mining (recurring session paths).
It surfaces issues at trace, session, and systemic levels by tracking user goals/constraints, progress and failure sequences, then grouping recurring high-impact failure modes rather than just clustering similar conversations.
The chances of a wrong fix is very low. However, even if it happens, the devs have the final say at "Merge the PR" stage. If it seems wrong, you can reply in the Github PR and selfship will improve upon it. If it's still going nowhere, the PR can be closed.
Yes, this is a very interesting use case. We are working with a fintech startup and a manufacturing enterprise to deploy this on premise. Would love to understand your use case better and work towards it. Please schedule a call with me at https://calendly.com/pr4n
Very cool. Curious if the results are better for certain domains vs others and if there's a sweet spot in terms of number of turns in the convos. What have you seen deploying this in the wild so far?
So far, our early design partners have been running this on (a) AI Based Ads creation and optimisation product [Multi Modal] (b) Customer Support for eCommerce [Text Based] and (c) our own AI trading system.
We’re actively looking for more use cases because we want to understand where the approach breaks down.
We haven't seen a clear domain specific sweet spot yet. What surprises us is that it becomes seems to get better as the complexity / turn count increases.
It adopts very well to multi turn sessions, scheduled runs and even trigger based Agentic systems that don't involve human in the loop.
We don’t yet have enough data to say something like “10–20 turns is optimal,” though. That’s one of the things we’re hoping to learn as we get it into more production systems.
What happens when Selfship proposes the wrong fix?
Great question. It usually never happens. Our system is based on an adoption of Domino (systematic error/slice discovery), AgentBoard (trajectory/progress evaluation), τ-bench (goal/outcome correctness), MAST (failure taxonomies) and process mining (recurring session paths).
It surfaces issues at trace, session, and systemic levels by tracking user goals/constraints, progress and failure sequences, then grouping recurring high-impact failure modes rather than just clustering similar conversations.
The chances of a wrong fix is very low. However, even if it happens, the devs have the final say at "Merge the PR" stage. If it seems wrong, you can reply in the Github PR and selfship will improve upon it. If it's still going nowhere, the PR can be closed.
A bad fix never lands.
Can I run this on-prem?
Yes, this is a very interesting use case. We are working with a fintech startup and a manufacturing enterprise to deploy this on premise. Would love to understand your use case better and work towards it. Please schedule a call with me at https://calendly.com/pr4n
[flagged]
[flagged]