Følg
Ilya Kostrikov
Ilya Kostrikov
OpenAI
Verifisert e-postadresse på openai.com - Startside
Tittel
Sitert av
Sitert av
År
Offline reinforcement learning with implicit q-learning
I Kostrikov, A Nair, S Levine
arXiv preprint arXiv:2110.06169, 2021
9562021
Image augmentation is all you need: Regularizing deep reinforcement learning from pixels
I Kostrikov*, D Yarats*, R Fergus
arXiv preprint arXiv:2004.13649, 2020
911*2020
Planet-photo geolocation with convolutional neural networks
T Weyand, I Kostrikov, J Philbin
Computer Vision–ECCV 2016: 14th European Conference, Amsterdam, The …, 2016
5472016
Improving sample efficiency in model-free reinforcement learning from images
D Yarats, A Zhang, I Kostrikov, B Amos, J Pineau, R Fergus
Proceedings of the aaai conference on artificial intelligence 35 (12), 10674 …, 2021
4912021
Intrinsic motivation and automatic curricula via asymmetric self-play
S Sukhbaatar, Z Lin, I Kostrikov, G Synnaeve, A Szlam, R Fergus
arXiv preprint arXiv:1703.05407, 2017
4502017
Discriminator-actor-critic: Addressing sample inefficiency and reward bias in adversarial imitation learning
I Kostrikov, KK Agrawal, D Dwibedi, S Levine, J Tompson
arXiv preprint arXiv:1809.02925, 2018
3592018
Offline Reinforcement Learning with Fisher Divergence Critic Regularization
I Kostrikov, J Tompson, R Fergus, O Nachum
arXiv preprint arXiv:2103.08050, 2021
3502021
Gpt-4o system card
A Hurst, A Lerer, AP Goucher, A Perelman, A Ramesh, A Clark, AJ Ostrow, ...
arXiv preprint arXiv:2410.21276, 2024
2932024
Training diffusion models with reinforcement learning
K Black, M Janner, Y Du, I Kostrikov, S Levine
arXiv preprint arXiv:2305.13301, 2023
2902023
Algaedice: Policy gradient from arbitrary experience
O Nachum, B Dai, I Kostrikov, Y Chow, L Li, D Schuurmans
arXiv preprint arXiv:1912.02074, 2019
2732019
Automatic data augmentation for generalization in deep reinforcement learning
R Raileanu, M Goldstein, D Yarats, I Kostrikov, R Fergus
arXiv preprint arXiv:2006.12862, 2020
244*2020
Pytorch implementations of reinforcement learning algorithms
I Kostrikov
GitHub repository: https://github.com/ikostrikov/pytorch-a2c-ppo-acktr-gail, 2018
2442018
Imitation learning via off-policy distribution matching
I Kostrikov, O Nachum, J Tompson
arXiv preprint arXiv:1912.05032, 2019
2302019
Rvs: What is essential for offline rl via supervised learning?
S Emmons, B Eysenbach, I Kostrikov, S Levine
arXiv preprint arXiv:2112.10751, 2021
2182021
An efficient convolutional network for human pose estimation.
U Rafi, B Leibe, J Gall, I Kostrikov
BMVC 1, 2, 2016
1792016
Efficient Online Reinforcement Learning with Offline Data
PJ Ball*, L Smith*, I Kostrikov*, S Levine
arXiv preprint arXiv:2302.02948, 2023
1742023
Idql: Implicit q-learning as an actor-critic method with diffusion policies
P Hansen-Estruch, I Kostrikov, M Janner, JG Kuba, S Levine
arXiv preprint arXiv:2304.10573, 2023
1382023
A walk in the park: Learning to walk in 20 minutes with model-free reinforcement learning
L Smith*, I Kostrikov*, S Levine
arXiv preprint arXiv:2208.07860, 2022
133*2022
Offline rl for natural language generation with implicit language q learning
C Snell, I Kostrikov, Y Su, M Yang, S Levine
arXiv preprint arXiv:2206.11871, 2022
1092022
Openai o1 system card
A Jaech, A Kalai, A Lerer, A Richardson, A El-Kishky, A Low, A Helyar, ...
arXiv preprint arXiv:2412.16720, 2024
1062024
Systemet kan ikke utføre handlingen. Prøv på nytt senere.
Artikler 1–20