Source code for our paper "Solving Discrete Logarithms in Smooth-Order Groups with CUDA" from SHARCS 2012. The cuda-11.5 branch will compile on CUDA 11.5 and now runs with two A100 GPUs, though is perhaps not tuned optimally.

Steven Engler 7ae53c058d Small change to Makefile.graham. hace 8 años
.gitignore 3a50b8f3a9 Added a '.gitignore' file. hace 8 años
COPYING e2973619c9 Commit cudadl-0.8 to git hace 14 años
Makefile 0ec3f2b5b8 Improved support for local Makefiles. hace 8 años
Makefile.graham 7ae53c058d Small change to Makefile.graham. hace 8 años
Makefile.ripple d90f89617c Added custom local Makefiles. hace 8 años
README 9fca8d29fd cudadl-0.9 hace 14 años
atomic_iostream.h 3a802c8e75 Improved the logging for dlrho. hace 8 años
controller.cc ecd76f9dec Improved scheduling. hace 8 años
controller.h ecd76f9dec Improved scheduling. hace 8 años
controller_main.cc ecd76f9dec Improved scheduling. hace 8 años
cudadl.h 1da839f0d8 Can save distinguished points to files. hace 8 años
desired_resources.cc ecd76f9dec Improved scheduling. hace 8 años
desired_resources.h ecd76f9dec Improved scheduling. hace 8 años
dlrho.cc ecd76f9dec Improved scheduling. hace 8 años
dpnode.cc 411725a245 Added warning when dpnode buffers don't empty. hace 8 años
dpnode.h 5b62cbb2c4 Split off the three main()s in preparation for MPI wrapper hace 14 años
dpnode_main.cc 5b62cbb2c4 Split off the three main()s in preparation for MPI wrapper hace 14 años
dpstream.cu 1da839f0d8 Can save distinguished points to files. hace 8 años
evutils.cc d8d46482b9 Hardcoded rules for hostname modifications. hace 8 años
evutils.h 411725a245 Added warning when dpnode buffers don't empty. hace 8 años
gen_N.cc 9fca8d29fd cudadl-0.9 hace 14 años
gen_all_N.sh 47faf09287 Improved the timing experiment script. hace 8 años
gencios_reg_20 771b13c42c Use macros to include the inline assembly instead of inline functions hace 14 años
logparse.pl 9fa203afbe Add a program to parse the logfile for some interesting statistics hace 14 años
mpi.cc 904cf8667e MPI version redirects some output to files. hace 8 años
parrhoasm.cu 5f32eec394 Can set n{threads/blocks} using a macro. hace 8 años
run_timing_experiment.sh 47faf09287 Improved the timing experiment script. hace 8 años
slurm-jobscript.sh e1af1cc644 Added scripts to use cudadl on Graham. hace 8 años
slurm-wrapper.sh dd03a46555 Fixed bug in Slurm wrapper script. hace 8 años
subproblem.h e5decaf5df Communicate the dpfreq to the workers correctly hace 14 años
summarize_data_multiprocess.py ce3141f7e1 Added scripts for collecting timing data. hace 8 años
worker.cc f619c464d8 Small bugfix. hace 8 años
worker.h 0e11651c7d The worker's 'cuda_dev_id' argument selects a GPU. hace 8 años
worker_main.cc 0e11651c7d The worker's 'cuda_dev_id' argument selects a GPU. hace 8 años

README

cudadl-0.9
21 Mar 2012
Ryan Henry and Ian Goldberg
{rhenry,iang}@cs.uwaterloo.ca
http://crysp.uwaterloo.ca/software/

This package contains the source code to our CUDA implementation of
van Oorschot and Wiener's parallel version of the Pollard rho discrete
log algorithm. It is intended for use on 1536-bit moduli that are
RSA numbers with smooth totient; that is, the modulus N=pq, where p and q
are 768-bit primes, and the prime factors of p-1 and q-1 are all
distinct and less then B, for a parameter B. [The value 1536 is
hardcoded as "WORDS = 24" (24*32*2 = 1536) in the Makefile; it is easy
to change this value and recompile if desired.] Note that this means
the totient of N = \phi(N) = (p-1)(q-1) has all prime factors less than
B; that is, \phi(n) is "B-smooth".

Usage:

1. Build the software. You'll need:

NTL
GMP
NVIDIA CUDA Toolkit 3.1
2 M2050 (or other compute capability level 2.0) CUDA cards
[If you have more or just 1, you'll need to modify dlrho.cc,
unfortunately.]

Hopefully just typing "make" should work. It will build gen_N and
dlrho.

2. Create the modulus N as, for example, a 1536-bit RSA number whose
totient is 2^50-smooth:

./gen_N 1536 50 > N

3. Generate a DL problem mod N and solve it:

./dlrho < N

This software is described in "Solving Discrete Logarithms in
Smooth-Order Groups with CUDA", CACR technical report 2012-02,
http://www.cacr.math.uwaterloo.ca/techreports/2012/cacr2012-02.pdf

This program is covered under version 3 of the GNU General Public
Licence; see the file COPYING for more information.

Changelog:

0.9 (21 Mar 2012)
Extend the code to handle smoothness levels (B) larger than 2^60. Now
we can handle up to 2^92. We have successfully run a test with
B = 2^80.

0.8 (23 Jan 2012)
Initial public release